3 ms·
> Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen We're not reading the same numbers I think. Compared to
by TacticalCoder 6mo ago
> Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen
We're not reading the same numbers I think. Compared to Opus 4.6, it's a big jump nearly in every single bench GP posted. They're "only" catching up to Google's Gemini on GPQA and MMMLU but they're still beating their own Opus 4.6 results on these two.
This sounds like a much better model than Opus 4.6.
- ninjagoo 6mo ago> We're not reading the same numbers I think. We must not be. That's why I listed out the ones where it is barely competitive from @babelfish's table, which itself is extracted from Pg 186 & 187 of the System Card, which has the comparison with Opus 4.6, GPT 5.4 and Gemini 3.1 Pro. Sure, it may be better than Opus 4.6 on some of those, but barely achieves a small increase over GPT-5.4 on the ones I called out.
- nimchimpsky 6mo agobarely competitive ? Mythos column is the first column. You are the only person with this take on hackernews, everyone else "this is a massive a jump". Fwiwi, the data you list shows the biggest jump I remember for mythos
- devmor 6mo agoThe biggest jump in the numbers they quoted is 6%. Please look at the columns OTHER than Opus as well.
- josephg 6mo ago> Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) > Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% > USAMO: 97.6% / 42.3% / 95.2% / 74.4% > The biggest jump in the numbers they quoted is 6%. Just in the numbers you quoted, thats a 16.6% jump in terminal-bench and a 55.3% absolute increase in USAMO over their previous Opus 4.6 model.
- devmor 6mo ago[flagged]
- dang 6mo agoCan you please stop posting comments with personal swipes in them? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- nimchimpsky 6mo ago[dead]
- nl 6mo agoIt's higher than all other models except vs Gemini 3.1 Pro on MMMLU
- DroneBetter 6mo ago[flagged]
- dang 6mo agoPlease don't respond to a bad comment by breaking the site guidelines yourself. That only makes things worse. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- nl 6mo ago> barely competitive It's higher than all other models except vs Gemini 3.1 Pro on MMMLU MMMLU is generally thought to be maxed out - as it it might not be possible to score higher than those scores. > Overall, they estimated that 6.5% of questions in MMLU contained an error, suggesting the maximum attainable score was significantly below 100%[1] Other models get close on GPQA Diamond, but it wouldn't be surprising to anyone if the max possible on that was around the 95% the top models are scoring. [1] https://en.wikipedia.org/wiki/MMLU https://en.wikipedia.org/wiki/MMLU
- lostmsu 6mo agoYou are reading the percentages wrong. Because 100% is maximum, you should be looking at error rates instead. GPT has 25% on Terminal Bench and the new model has 18%, almost 1.4x reduction.