4 ms·
Might still be fine. The most recent crop of vLLMs proactively use whichever programs are available on the system (e.g. ImageMagick or PIL) to "zoom in" by crop
by johndough 1mo ago
Might still be fine. The most recent crop of vLLMs proactively use whichever programs are available on the system (e.g. ImageMagick or PIL) to "zoom in" by cropping subimages if they can't quite make out the details.
- knollimar 1mo agoDownsizing a higher res image to lower res means the zoom will be blurry.
- andai 1mo agoThey process the original image file with Python on the local device. (And I've seen the web chats do this with their "computer use" features too.) The really wild one is even blind models will do this and they'll try to run stats on the pixels to figure out what it looks like... the even wilder thing is that it kind of works!
- knollimar 1mo agoIf the API accepts only 800 by 800, the aegument youre making is "fix it in the harness". I don't think the n by n subgrid fixes this the way most harnesses do, as it'll fail to count things if you have more overlap and fail relatiomships if you have less
- tjoff 1mo agoThat seems weirdly specific? And if you are counting things it should be trivial to note the position of your items and not double-count them, no?
- adastra22 1mo agoThey’re not talking about zooming, hence the quotes.
- knollimar 1mo agoIf the harness does it that's just like saying "please use a workaround". You'll lose fidelity and LLMs will lose the ability to count things or maintain relationships for schematics, etc
- johndough 1mo agoLLMs read images by splitting them up into e.g. 16x16 patches, which are then converted to embedding vectors and fed to the LLM, so from a technical point of view, feeding a big image as many 20x20 patches all at once is not too different from cropping subimages from the image, splitting those subimages into patches and feeding them to the LLM. Of course, the LLM has to be trained to understand that those images belong together, but it can be done.
- knollimar 1mo agoYou say "it can be trained" but they fail at counting in my use cases, let alone maintaining symbolic relationships. Can you point me at one that can understand a detailed block diagram? Frontier is fine, soliciting recommendations
- johndough 1mo agoFor counting, there are specialized counting models, e.g. https://huggingface.co/spaces/MengqiLei/count-anything-demo https://huggingface.co/spaces/MengqiLei/count-anything-demo I tried to parse hand-drawn ER diagrams in the past and did not have much success with any model, frontier or otherwise. If you have annotated data, I'd recommend finetuning a recent (dense) VLM, but don't expect 100% accuracy. https://unsloth.ai/docs/basics/vision-fine-tuning https://unsloth.ai/docs/basics/vision-fine-tuning
- knollimar 1mo agogood advice; tbh I'm really trying to use counting as a proxy for "understand symbol, flag annotation next to it, associate annotation with symbol and store as object in location". Ideally, look at second diagram and see something similar in the same location and understand it's the same physical object.
- johndough 1mo agoThe order is: LLM issues tool call to read high res image -> harness sends high res image to server -> server downsizes it to 800x800 (blurry) -> LLM issues bash command (e.g. `convert`) to crop a small subimage (e.g. 600x600) from the high res image -> LLM issues tool call to read subimage -> harness sends subimage to server -> server does not resize the subimage because it is small already, so it is not blurry when finally ingested by the LLM
- deleted 1mo ago[deleted]
- knollimar 1mo agoThen you have a separate issue where the LLM can't piece together 9 subimages well.
- nwienert 1mo agoInsert the famous Louie CK phones on planes bit. Sand is thinking, working, coding and now looking for you, for pennies an hour... but "oof" it's not good enough.