3 ms·
Actually - do they do this in LLM benchmarks? As a measure of overconfidence/confabulation? Seems immediately applicable.
by andyferris 8mo ago
Actually - do they do this in LLM benchmarks? As a measure of overconfidence/confabulation? Seems immediately applicable.
- impossiblefork 8mo agoI don't think it's a common thing in any public LLM benchmarks or in any standard QA datasets. Maybe in internal stuff at AI firms.