2 ms·
Do you publish your benchmark somewhere? What's the best vision (captioning) AI you've come across?
by addandsubtract 1mo ago
Do you publish your benchmark somewhere? What's the best vision (captioning) AI you've come across?
- jerkstate 1mo agoMy benchmark is tiny compared to WorldVQA or FG-BMK, which are available, so I'd point you in that direction if you're interested in a VLM benchmark. My use-case isn't exactly captioning as in "what is in this image?" -> caption, I am using the VLM to validate captions, as in "is this an image of [supposed subject]?" - my ranking of models I've benchmarked is gemini-3.7-flash > seed-2.1-turbo > gpt-5.6-luna > qwen-3.7-plus > qwen-3.7-flash. Gemini is almost perfect on my test dataset, only failing on some esoteric pop-culture minor celebrities and being over-specific in some cases (i.e. Q: is this [common name of fruit]? A: that's a [latin species name of fruit], not a [common name of fruit]; false). However, gemini-3.7-flash is only in my test list because openrouter has it on 75% introductory discount; otherwise it would be about 4x more expensive than seed.