4 ms·
On this topic, SimpleQA benchmark has a component measuring hallucination rate vs ”know” vs ”don’t know”. OpenAI models have often been more troubled than the r
by jug 2y ago
On this topic, SimpleQA benchmark has a component measuring hallucination rate vs ”know” vs ”don’t know”. OpenAI models have often been more troubled than the rest. See also, from the paper: https://imgur.com/7NDZ0ON https://imgur.com/7NDZ0ON (you want a low ”Incorrect” score as it’s an attempted answer, but wrong)
I wish hallucination benchmarks were far more popular.