4 ms·
This model is actually pretty bad. Sure, it can do things like solve math equations from an image, but the vision part of that is basic OCR. In terms of actual
by Jackson__ 2y ago
This model is actually pretty bad. Sure, it can do things like solve math equations from an image, but the vision part of that is basic OCR. In terms of actual vision capabilities, i.e. understanding dense images correctly, these models all fail the same.
Over the past year, researchers have been benchmark chasing, without caring about the actual abilities of these models. This is especially damning in the vision space, where most "vision" benchmarks consist entirely of either of leading questions that trivialize the image understanding part, or are straight up A PNG version of a regular LLM benchmark, once again reducing the importance of vision down to being able to OCR.
Due to this, it appears the main take-away for researchers in the field has been "Vision Ability improves with LLM size." and I wish I was joking. This willful misunderstanding is reflected in their architectural choices. 99% of current VLM rely on some flavor of CLIP models to interpret images, with very little attempts made to improve beyond that as it does not show improvements on current benchmarks.
Let's take one classic example picture, I like to feed these models just to see if maybe I was wrong about all this. The task is simple. Describe the image. https://preview.redd.it/hqu05vlmzu8e1.png?width=1516&format=png&auto=webp&s=db5ec72882b04476753ab7285f0b50c5e081ee33 https://preview.redd.it/hqu05vlmzu8e1.png?width=1516&format=...
It utterly fails, and hallucinates a ridiculous amount while doing so. There is honestly about as much wrong about the description as there is right. No current VLM open source or closed can properly find who is holding the candy bucket in the front. None. Not Claude, ChatGPT, Gemini, Llama3.2, Qwen, or anyone else.
And it is all because benchmarks are king, and innovation is dead.
- letmevoteplease 2y agoChatGPT, Claude and Gemini get everything correct except who is holding the bucket. Even the QvQ attempt in your screenshot would have seemed like complete magic a couple years ago.
- jsheard 2y agoLike how LLMs tend to get tripped up by questions that are phrased like well-known riddles, it turns out you can trip up vision models with images that resemble well-known optical illusions. Overfitting strikes again. https://x.com/kosa12matyas/status/1871256745403953572 https://x.com/kosa12matyas/status/1871256745403953572
- Jackson__ 2y agoIn essence, that is what HallusionBench[0] does. While I think it is an improvement over other vision benchmarks, it still falls short in terms of quantifying actual vision capabilities. More than anything, it seems like a way to detect whether the model was over trained on these riddles. [0] https://github.com/tianyi-lab/HallusionBench https://github.com/tianyi-lab/HallusionBench
- Kerbonut 2y agoBenchmarks are useful to know where MLMs (multimodal language models) are deficient. Without an automated way to test improvements, improvements are rarely made. I don't think it's really about innovation being dead as much as like you said, benchmarks are king. So I would say if you want things to improve in that area, create a benchmark and it will eventually get solved.