4 ms·
It's really interesting that there's a huge performance discrepancy between these SOTA models. In the Olympic logo example, GPT-4o is below the baseline accurac
by _vaporwave_ 2y ago
It's really interesting that there's a huge performance discrepancy between these SOTA models. In the Olympic logo example, GPT-4o is below the baseline accuracy of 20% (worse than randomly guessing) while Sonnet-3.5 was correct ~76% of the time.
Does anyone have any technical insight or intuition as to why this large variation exists?
- ec109685 2y agoThe question wasn’t “yes or no” but instead required an exact number: https://huggingface.co/datasets/XAI/vlmsareblind/viewer/default/train https://huggingface.co/datasets/XAI/vlmsareblind/viewer/defa... Playing around with GPT-4o, it knows enough to make a copy of an image that is reasonable but it still can’t answer the questions. ChatGPT went down a rabbit hole of trying to write python code, but it took lots of prompting for it to notice its mistake when solving one of the intersecting line questions.