2 ms·
92% suggests a harder benchmark is needed, so it's difficult judge. Especially when a lot of "high scoring" models produce cogent results with a high level of h
by joshvm 2y ago
92% suggests a harder benchmark is needed, so it's difficult judge. Especially when a lot of "high scoring" models produce cogent results with a high level of hallucination (eg Llama 3 is chatty, confident and quite often wrong for me).
At that level of performance you're probably in the realm of hard edge cases with ambiguous ground truth.