3 ms·
We should be able to find something that was missed! Did you let any LLMs search this?
by archleaf 5mo ago
We should be able to find something that was missed!
Did you let any LLMs search this?
- keepamovin 5mo agoYeah, actually I think that’s really smart. Because after you convert everything to JPEG everything is just an image that you can ask LLMs to look at. Unfortunately, I don’t have the experience with local models, but if someone wants to point me in like the right direction or send me an email to collab.
- vunderba 5mo agoThere are actually a few capable VL models out there that can run on even modest hardware. If you want to keep things simple and process everything locally, I’d recommend something like Qwen3 VL [1]. It’s not the fastest model, but you can just let it chew through the docs over a weekend. In my experience, it takes about 15 to 30 seconds per image, but the quality of the results is quite good if a bit verbose [2]. [1] - https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-FP8 https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-FP8 [2] - https://mordenstar.com/other/vlm-xkcd https://mordenstar.com/other/vlm-xkcd