3 ms·
I don't like this paper for the following reasons: - The language is unnecessarily scathing - They repeatedly show data where the models are getting things ri
by pjs_ 2y ago
I don't like this paper for the following reasons:
- The language is unnecessarily scathing
- They repeatedly show data where the models are getting things right 70, 80, 90% of the time, and then show a list of what they call "qualitative samples" (what does "qualitative" mean? "cherry-picked"?) which look very bad. But it got the answer right 70/80/90% of the time! That's hardly "blind"...
- Various of the tasks hinge on the distinction between two objects "exactly touching" vs. "very nearly touching" vs. "very slightly overlapping", a problem which (i) is hard for humans and (ii) is particularly (presumably deliberately) sensitive to resolution/precision, where we should not be surprised that models fail
- The main fish-shaped example given in task 1 seems genuinely ambiguous to me - do the lines "intersect" once or twice? The tail of the fish clearly has a crossing, but the nose of the fish seems a bit fishy to me... is that really an intersection?
- AFAIC deranged skepticism is just as bad as deranged hype, the framing here is at risk of appealing to the former
It's absolutely fair to make the point that these models are not perfect, fail a bunch of the time, and to point out the edge cases where they suck. That moves the field forwards. But the hyperbole (as pointed out by another commenter) is very annoying.
- neuronet 2y agoTo be fair, the paper has an emoji in the _title_, so I wouldn't read it as a particularly particularly serious academic study as much as the equivalent of the Gawker of AI research. It is a "gotcha" paper that exploits some blind spots (sorry) that will easily be patched up with a few batches of training. I do think it highlights the lack of AGI in these things, which some people lacking situational awareness might need to see.
- numeri 2y agoI'm also confused about some of the figures' captions, which don't seem to match the results: - "Only Sonnet-3.5 can count the squares in a majority of the images", but Sonnet-3, Gemini-1.5 and Sonnet-3.5 all have accuracy of >50% - "Sonnet-3.5 tends to conservatively answer "No" regardless of the actual distance between the two circles.", but it somehow gets 91% accuracy? That doesn't sound like it tends to answer "No" regardless of distance.
- schneehertz 2y agoI am not sure where their experimental data came from. I tested it on GPT-4o using the prompt and images they provided, and the success rate was quite high, with significant differences from the results they provided.
- ec109685 2y agoTheir examples are here: https://huggingface.co/datasets/XAI/vlmsareblind/viewer/default/train https://huggingface.co/datasets/XAI/vlmsareblind/viewer/defa... ChatGPT whiffs completely on very obvious images.