4 ms·
Meta got caught _first_.
by etamponi 2y ago
Meta got caught _first_.
- sumeno 2y agoNot even first, OpenAI got caught a while back
- Mond_ 2y agoDo you have a source for this? That's interesting (if true).
- sumeno 2y agoThey got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it https://techcrunch.com/2025/01/19/ai-benchmarking-organization-criticized-for-waiting-to-disclose-funding-from-openai/ https://techcrunch.com/2025/01/19/ai-benchmarking-organizati...
- tananaev 2y agoI don't see anything in the article about being caught. Maybe I missed something?
- tedsanders 2y ago[flagged]
- suddenlybananas 2y ago>They got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it Is anything here actually false or do you not like the conclusions that people may draw from it?
- tedsanders 1y agoThe false statement is “OpenAI got caught [gaming FrontierMath] a while back.” From a primary source: "OpenAI did not use FrontierMath data to guide the development of o1 or o3, at all.... we only downloaded FrontierMath for our evals long after the training data was frozen, and only looked at o3 FrontierMath results after the final announcement checkpoint was already picked." https://x.com/__nmca__/status/1882563755806281986 https://x.com/__nmca__/status/1882563755806281986
- suddenlybananas 1y agoWell, if we trust you. But you had the extra dataset and no one else did.
- tucnak 2y agoWhy are you being disingenuous? Simply having access to the eval in question is already enough for your synthetics guys to match the distribution, and of course you don't contaminate directly on train, that would be stupid, and you would get caught, but if it does inform the reward, the result is the same. You _should_ quit, but you wouldn't because you'd already convinced yourself you're doing RL God's work, not sleight of hand. > If an MMLU question asks about an abstract algebra proof, is it cheating to have trained on papers about abstract algebra? This kind of disingenuous bullshit is exactly why people call you cheaters. > Generally, I don’t think anyone here is cheating and I think we’re relatively diligent with our evals. You guys should follow Apple's cult guidelines: never stand out. Think different
- tedsanders 1y agoFrom a primary source: "OpenAI did not use FrontierMath data to guide the development of o1 or o3, at all.... we only downloaded FrontierMath for our evals long after the training data was frozen, and only looked at o3 FrontierMath results after the final announcement checkpoint was already picked." https://x.com/__nmca__/status/1882563755806281986 https://x.com/__nmca__/status/1882563755806281986
- FridgeSeal 2y ago[flagged]
- tomhow 1y agoBe kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- sumeno 2y ago[flagged]
- tomhow 1y agoBe kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- sumeno 1y agoWhat about the accusation that I was spreading false rumors? That's far more unkind
- tomhow 1y agoThe guidelines apply to everyone. I didn't like that part of their comment either, and the comment would have been better without it, but they were put on the defensive by what had come before and they went on to explain their position. Your phrase "got caught with your hand in the cookie jar" is a swipe that the thread could also have done without. It's no big deal, and it applies to everyone on the subthread; we want everyone to avoid barbs like that on HN.
- tedsanders 1y agoThe incorrect part is “OpenAI got caught [gaming FrontierMath] a while back.” From a primary source: "OpenAI did not use FrontierMath data to guide the development of o1 or o3, at all.... we only downloaded FrontierMath for our evals long after the training data was frozen, and only looked at o3 FrontierMath results after the final announcement checkpoint was already picked." https://x.com/__nmca__/status/1882563755806281986 https://x.com/__nmca__/status/1882563755806281986
- tomhow 1y agoYour comment would be better without the personal swipe in the first sentence. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- wongarsu 2y agoIt happens with basically all papers on all topics. Benchmarks are useful when they are first introduced and used to measure things that were released before the benchmark. After that their usefulness rapidly declines.
- hooloovoo_zoo 2y agoPeople have been gaming ML benchmarks as long as there have been ML benchmarks. That's why it's better to see if other researchers are incorporating a technique into their actual models rather than 'is this paper the bold entry in a benchmark table'. But it takes longer.
- jandrese 2y agoWhen a measure becomes a target it is no longer a good measure. These ML benchmarks were never going to last very long. There is far too much pressure to game them, even unintentionally.
- mkolodny 2y ago“Got caught” is a misleading way to present what happened. According to the article, Meta publicly stated, right below the benchmark comparison, that the version of Llama on LMArena was the experimental chat version: > According to Meta’s own materials, it deployed an “experimental chat version” of Maverick to LMArena that was specifically “optimized for conversationality” The AI benchmark in question, LMArena, compares Llama 4 experimental to closed models like ChatGPT 4o latest, and Llama performs better (https://lmarena.ai/?leaderboard https://lmarena.ai/?leaderboard).