4 ms·
It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before Oc
by jug 2y ago
It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.
- tmpz22 2y agoAre the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?
- llm_trw 2y agoIn general yes, bench mark pollution is a big problem and why only dynamic benchmarks matter.
- brookst 2y agoThis is true, but how would pollution work for a benchmark designed to test hallucinations?
- llm_trw 2y agoA dataset of labelled answers that are hallucinations and not hallucinations are published based on the benchmark as part of a paper. People _seriously_ underestimate just how much stuff is online and how much impact it can have on training.
- sumeno 2y agoThey've been caught in the past getting benchmark data under the table, if they got caught once they're probably doing it even more
- refulgentis 2y agoNo, they haven't.
- freehorse 2y agoThey actually have [0]. They were revealed to have had access to the (majority of the) frontierMath problemset while everybody thought the problemset was confidential, and published benchmarks for their o3 models on the presumption that they didn't. I mean one is free to trust their "verbal agreement" that they did not train their models on that, but access they did have and it was not revealed until much later. [0] https://the-decoder.com/openai-quietly-funded-independent-math-benchmark-before-setting-record-with-o3/ https://the-decoder.com/openai-quietly-funded-independent-ma...
- brookst 2y agoCurious you left out Frontier Math’s statement that they provided 300 questions plus answers, and another holdback set of 50 questions without answers, to allay this concern. [0] We can assume they’re lying too but at some point “everyone’s bad because they’re lying, which we know because they’re bad” gets a little tired. 0. https://epoch.ai/blog/openai-and-frontiermath https://epoch.ai/blog/openai-and-frontiermath
- freehorse 2y ago1. I said the majority of the problems, and the article I linked also mentioned this. Nothing “curious” really, but if you thought this additional source adds sth more, thanks for adding it here. 2. We know that “open”ai is bad, for many reasons, but this is irrelevant. I want processes themselves to not depend on the goodwill of a corporation to give intended results. I do not trust benchmarks that first presented themselves secret and then revealed they were not, regardless if the product benchmarked was from a company I otherwise trust or not.
- brookst 2y agoFair enough. It’s hard for me to imagine being so offended as the way they screwed up disclosure that I’d reject empirical data, but I get that it’s a touchy subject.
- ipaddr 2y agoBenchmarks are not real so 2% is meaningless.
- fn-mote 2y agoOf course not. The point is that the cost difference between the two things being compared is huge, right? Same performance, but not the same cost.
- random_cynic 2y ago[dead]
- refulgentis 2y agoIt's not SimpleQA...
- Gazoche 2y agoI wonder how it's even possible to evaluate this kind of thing without data leakage. Correct answers to specific, factual questions are only possible if the model has seen those answers in the training data, so how reliable can the benchmark be if the test dataset is contaminated with training data? Or is the assumption that the training set is so big it doesn't matter?