4 ms·
Are the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?
by tmpz22 2y ago
Are the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?
- llm_trw 2y agoIn general yes, bench mark pollution is a big problem and why only dynamic benchmarks matter.
- brookst 2y agoThis is true, but how would pollution work for a benchmark designed to test hallucinations?
- llm_trw 2y agoA dataset of labelled answers that are hallucinations and not hallucinations are published based on the benchmark as part of a paper. People _seriously_ underestimate just how much stuff is online and how much impact it can have on training.
- sumeno 2y agoThey've been caught in the past getting benchmark data under the table, if they got caught once they're probably doing it even more
- refulgentis 2y agoNo, they haven't.
- freehorse 2y agoThey actually have [0]. They were revealed to have had access to the (majority of the) frontierMath problemset while everybody thought the problemset was confidential, and published benchmarks for their o3 models on the presumption that they didn't. I mean one is free to trust their "verbal agreement" that they did not train their models on that, but access they did have and it was not revealed until much later. [0] https://the-decoder.com/openai-quietly-funded-independent-math-benchmark-before-setting-record-with-o3/ https://the-decoder.com/openai-quietly-funded-independent-ma...
- brookst 2y agoCurious you left out Frontier Math’s statement that they provided 300 questions plus answers, and another holdback set of 50 questions without answers, to allay this concern. [0] We can assume they’re lying too but at some point “everyone’s bad because they’re lying, which we know because they’re bad” gets a little tired. 0. https://epoch.ai/blog/openai-and-frontiermath https://epoch.ai/blog/openai-and-frontiermath
- freehorse 2y ago1. I said the majority of the problems, and the article I linked also mentioned this. Nothing “curious” really, but if you thought this additional source adds sth more, thanks for adding it here. 2. We know that “open”ai is bad, for many reasons, but this is irrelevant. I want processes themselves to not depend on the goodwill of a corporation to give intended results. I do not trust benchmarks that first presented themselves secret and then revealed they were not, regardless if the product benchmarked was from a company I otherwise trust or not.
- brookst 2y agoFair enough. It’s hard for me to imagine being so offended as the way they screwed up disclosure that I’d reject empirical data, but I get that it’s a touchy subject.
- 542354234235 2y agoWhen the data is secret and unavailable to the company before the test, it doesn’t rely on me trusting the company. When the data is not secret and is available to the company, I have to trust that the company did not use that prior knowledge to their advantage. When the company lies and says it did not have access, then later admits that it did have access, is means the data is less trustworthy from my outsider perspective. I don’t think “offense” is a factor at all. If a scientific paper comes out with “empirical data”, I will still look at the conflicts of interest section. If there are no conflicts of interest listed, but then it is found out that there are multiple conflicts of interest, but the authors promise that while they did not disclose them, they also did not affect the paper, I would be more skeptical. I am not “offended”. I am not “rejecting” the data, but I am taking those factors into account when determining how confident I can be in the validity of the data.