4 ms·
Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883 https://x.com/quietnning/status/2080786711861407883
by KaoruAoiShiho 2mo ago
Appears to be benchmaxxing
https://x.com/quietnning/status/2080786711861407883 https://x.com/quietnning/status/2080786711861407883
- throwa356262 2mo ago"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??
- pierrefermat1 2mo agoWhat seems to implied is that some of his hold out testing suite includes simple/common tests that are out in the wild, and for those opus went straight to a memorised solution .
- zamadatix 2mo agoSimple/common tests is not an explanation for why only now Opus 5 is the only model encoding the answers like this. Something like the holdout test suite being leaked or Anthropic cheating (e.g. 'accidentally' including previous hold out run data in Opus 5 training) makes a much stronger fit.
- stogot 2mo agoIt may read information about the benchmark, such as on blog post, or Twitter feeds (example OP) without active cheat
- modeless 2mo agoThey state the puzzle is "Witness-like" which I assume means that it follows the rules from the well-known puzzle game "The Witness" which Opus definitely knows.
- rbuccigrossi 2mo agoNo, I believe you are misunderstanding the quote. Each “question” in ARC-AGI-3 is a game that has hidden rules that you can understand if you look at the game board long enough. This quote means that Opus 5 is looking at the game board, figuring out the rules, and writing out the rules before it makes a single move. You can do the same thing if you go to the ARC-AGI-3 website and try some of the games.
- jnwatson 2mo agoI was just thinking they need to mark each model per benchmark as "model released before the benchmark was released" and "model released after the benchmark was released".
- andrepd 2mo agoI'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.
- asdfologist 2mo agoMeh, I doubt it was intentional. Deliberate benchmaxxing is incredibly damaging to credibility once it's discovered (see what happened to Meta with LlaMa 4). It's more likely that the training data was contaminated with the benchmark data.
- mupuff1234 2mo agoYou really think they saw the jump in arc-agi-3 (which they reported in their official card), and didn't even bother to check? They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .
- johnfn 2mo agoHow do you propose they check for something like this? They can't exactly ctrl-f the model weights for "Arc-AGI".
- phoghed 2mo ago> they can’t possibly know or find out what was in the training data doesn’t appear to be a very strong argument
- turing_curious 2mo ago[dead]
- vickychijwani 2mo agoAnthropic expends tons of compute and effort on understanding internal model states [1]; this kind of thing is right up their alley. [1]: Recent example: https://www.anthropic.com/research/global-workspace https://www.anthropic.com/research/global-workspace
- nozzlegear 2mo agohttps://xcancel.com/quietnning/status/2080786711861407883 https://xcancel.com/quietnning/status/2080786711861407883
- jchw 2mo agoI have been claiming that I don't think Chinese AI companies are benchmaxxing harder than American AI companies, which has gotten mixed reception: sometimes people agree, sometimes they disagree. It seems I was wrong. American AI companies might actually be benchmaxxing harder.
- gertlabs 2mo agoWe run an evaluation that is designed to be less vulnerable to benchmaxxing because there aren't correct solutions; agents are interacting in the same environment as other agents. And it's private, and our public benchmark is not well known enough for anyone to probably care to benchmax us yet. So I think it's pretty indicative of true relative aptitude. All models have probably memorized significant swaths of solution sets for popular benchmarks at this point, either accidentally or intentionally, so it's all relative at this point. However, in our experience, Chinese models do benchmax harder. This is also consistent with interacting with Chinese labs soliciting data/environments, who literally asked us for datasets and tasks modeled around and formatted like popular benchmarks. Opus 5 will be uploaded tomorrow, but we already have the tests locally and it is truly as capable as Fable, but at 81% of the real cost. (And from subjective usage, it has a very different personality) Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- jchw 2mo agoBenchmaxxing via memorization is boring and doesn't fool anyone for too long. It works, but then new benchmarks test old models and the real results fall in line. Benchmaxxing by focusing on specific types of things that benchmarks test on, while still not improving intelligence or capability in the general case? Not only is it blatantly obvious that all AI labs do this, but it's not even obvious how you would go about it any other way. Now I am not really specifically accusing Anthropic of anything here, I'm just saying their behavior is suspicious. Since you tested Fable, they wouldn't even have to lie to have optimized for your specific benchmarks, since they absolutely had permission to read your sessions if they wanted to. But obviously, that's only the situation if we take them at their word. Personally I would be a bit surprised if they just flat out were lying and secretly retaining data they say they are not, but not that surprised. The penalties for doing this are probably worth the rewards if it keeps them super far ahead in the benchmarks for years without anyone catching on. (In actuality though, even if they really were trying to sneakily grab samples of benchmark tests via their Fable data retention rules, I don't really suspect there would've been very much time to optimize Opus 5 on it. So consider me bothered.)
- codewiththiha 2mo ago[dead]
- root-parent 2mo agoAnd being worst than previous model... "...The traces tell the why: (1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre. (2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one..."
- Zababa 2mo ago>That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning. They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.
- FairyFair_ 2mo agoHello, sorry. I know this is not relevant to what you are talking about. However I have been unable to contact you by any other means and have been trying to contact you for more than a month Assuming that you are "kas", please check the dulst discord. It is important.