2 ms·
In essence, that is what HallusionBench[0] does. While I think it is an improvement over other vision benchmarks, it still falls short in terms of quantifying a
by Jackson__ 2y ago
In essence, that is what HallusionBench[0] does. While I think it is an improvement over other vision benchmarks, it still falls short in terms of quantifying actual vision capabilities. More than anything, it seems like a way to detect whether the model was over trained on these riddles.
[0] https://github.com/tianyi-lab/HallusionBench https://github.com/tianyi-lab/HallusionBench