14 ms·
Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro:
by babelfish 6mo ago
Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro)
SWE-bench Verified: 93.9% / 80.8% / — / 80.6%
SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2%
SWE-bench Multilingual: 87.3% / 77.8% / — / —
SWE-bench Multimodal: 59.0% / 27.1% / — / —
Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5%
GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3%
MMMLU: 92.7% / 91.1% / — / 92.6–93.6%
USAMO: 97.6% / 42.3% / 95.2% / 74.4%
GraphWalks BFS 256K–1M: 80.0% / 38.7% / 21.4% / —
HLE (no tools): 56.8% / 40.0% / 39.8% / 44.4%
HLE (with tools): 64.7% / 53.1% / 52.1% / 51.4%
CharXiv (no tools): 86.1% / 61.5% / — / —
CharXiv (with tools): 93.2% / 78.9% / — / —
OSWorld: 79.6% / 72.7% / 75.0% / —
- pants2 6mo agoWe're gonna need some new benchmarks... ARC-AGI-3 might be the only remaining benchmark below 50%
- randomtoast 6mo agoHumanity's Last Exam (HLE) is already insanely difficult. It introduces 2,500 questions spanning mathematics, humanities, natural sciences, ancient languages, ... Here is an example question: https://i.redd.it/5jl000p9csee1.jpeg https://i.redd.it/5jl000p9csee1.jpeg No human could even score 5% on HLE.
- saberience 6mo agoI've never understood the point of things like HLE, it doesn't really prove or show anything since 99.99% of humans can't do a single question on this exam. That is, it's easy to make benchmarks which humans are bad at, humans are really bad at many things. Divide 123094382345234523452345111 by 0.1234243131324, guess what, humans would find that hard, computers easy. But it doesn't mean much. Humanity's last exam (HLE) couldn't be completed by most of humanity, the vast majority, so it doesn't really capture anything about humanity or mean much if a computer can do it.
- DroneBetter 6mo agothe point is that each question is something that a specialist in a field would be able to do, but deems challenging enough that the ability to solve it would imply significant general usefulness in that domain
- Leynos 6mo agoOpus 4.6 currently leads the remote labor index at 4.17. GPT-5.4 isn't measured on that one though: https://www.remotelabor.ai/ https://www.remotelabor.ai/ GPT 5.4 Pro leads Frontier Maths Tier 4 at 35%: https://epoch.ai/benchmarks/frontiermath-tier-4/ https://epoch.ai/benchmarks/frontiermath-tier-4/
- mbesto 6mo ago> We're gonna need some new benchmarks... You can't consistently benchmark something that is qualitative by nature. I'm struggling to understand how people don't understand this.
- sourcecodeplz 6mo agoHaven't seen a jump this large since I don't even know, years? Too bad they are not releasing it anytime soon (there is no need as they are still currently the leader).
- ru552 6mo agoThere's speculation that next Tuesday will be a big day for OpenAI and possibly GPT 6. Anthropic showed their hand today.
- enraged_camel 6mo agoThat does not sound very believable. Last time Anthropic released a flagship model, it was followed by GPT Codex literally that afternoon.
- cyanydeez 6mo agoYa'll know they're teaching to the test. I'll wait till someone devises a novel test that isn't contained in the datasets. Sure, they're still powerful.
- swalsh 6mo agoMy understanding is GPT 6 works via synaptic space reasoning... which I find terrifying. I hope if true, OpenAI does some safety testing on that, beyond what they normally do.
- notrealyme123 6mo agoThat's sounds really interesting. Do you have some hints where to read more?
- levocardia 6mo agoOh you mean literally the thing in AI2027 that gets everyone killed? Wonderful.
- whalesalad 6mo agoHonestly we are all sleeping on GPT-5.4. Particularly with the influx of Claude users recently (and increasingly unstable platform) Codex has been added to my rotation and it's surprising me.
- rafaelmn 6mo agoGPT is shit at writing code. It's not dumb - extra high thinking is really good at catching stuff - but it's like letting a smart junior into your codebase - ignore all the conventions, surrounding context, just slop all over the place to get it working. Claude is just a level above in terms of editing code.
- whalesalad 6mo agoThis has been my experience. With very very rigid constraints it does ok, but without them it will optimize expediency and getting it done at the expense of integrating with the broader system.
- ctoth 6mo agoMy favorite example of this from last night: Me: Let's figure out how to clone our company Wordpress theme in Hugo. Here're some tools you can use, here's a way to compare screenshots, iterate until 0% difference. Codex: Okay Boss! I did the thing! I couldn't get the CSS to match so I just took PNGs of the original site and put them in place! Matches 100%!
- leobuskin 6mo agoAnd as a bonus: GPT is slow. I’m doing a lot of RE (IDA Pro + MCP), even when 5.4 gives a little bit better guesses (rarely, but happens) - it takes x2-x4 longer. So, it’s just easier to reiterate with Opus
- blazespin 6mo agoYeah, need some good RE benchmarks for the LLMs. :) RE is very interesting problem. A lot more that SWE can be RE'd. I've found the LLMs are reluctant to assist, though you can workaround.
- simianwords 6mo agoThe real part is SWE-bench Verified since there is no way to overfit. That's the only one we can believe.
- ollin 6mo agoMy impression was entirely the opposite; the unsolved subset of SWE-bench verified problems are memorizable (solutions are pulled from public GitHub repos) and the evaluators are often so brittle or disconnected from the problem statement that the only way to pass is to regurgitate a memorized solution. OpenAI had a whole post about this, where they recommended switching to SWE-bench Pro as a better (but still imperfect) benchmark: https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/ https://openai.com/index/why-we-no-longer-evaluate-swe-bench... > We audited a 27.6% subset of the dataset that models often failed to solve and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions > SWE-bench problems are sourced from open-source repositories many model providers use for training purposes. In our analysis we found that all frontier models we tested were able to reproduce the original, human-written bug fix > improvements on SWE-bench Verified no longer reflect meaningful improvements in models’ real-world software development abilities. Instead, they increasingly reflect how much the model was exposed to the benchmark at training time > We’re building new, uncontaminated evaluations to better track coding capabilities, and we think this is an important area to focus on for the wider research community. Until we have those, OpenAI recommends reporting results for SWE-bench Pro.
- simianwords 6mo agoI stand corrected.
- simianwords 6mo ago> My impression was entirely the opposite; the unsolved subset of SWE-bench verified problems are memorizable (solutions are pulled from public GitHub repos) and the evaluators are often so brittle or disconnected from the problem statement that the only way to pass is to regurgitate a memorized solution. Anthropic accounts for this >To detect memorization, we use a Claude-based auditor that compares each model-generated patch against the gold patch and assigns a [0, 1] memorization probability. The auditor weighs concrete signals—verbatim code reproduction when alternative approaches exist, distinctive comment text matching ground truth, and more—and is instructed to discount overlap that any competent solver would produce given the problem constraints.
- WarmWash 6mo agoAre these fair comparisons? It seems like mythos is going to be like a 5.4 ultra or Gemini Deepthink tier model, where access is limited and token usage per query is totally off the charts.
- mulmboy 6mo agoThere are a few hints in the doc around this > Importantly, we find that when used in an interactive, synchronous, “hands-on-keyboard” pattern, the benefits of the model were less clear. When used in this fashion, some users perceived Mythos Preview as too slow and did not realize as much value. Autonomous, long-running agent harnesses better elicited the model’s coding capabilities. (p201) ^^ From the surrounding context, this could just be because the model tends to do a lot of work in the background which naturally takes time. > Terminal-Bench 2.0 timeouts get quite restrictive at times, especially with thinking models, which risks hiding real capabilities jumps behind seemingly uncorrelated confounders like sampling speed. Moreover, some Terminal-Bench 2.0 tasks have ambiguities and limited resource specs that don’t properly allow agents to explore the full solution space — both being currently addressed by the maintainers in the 2.1 update. To exclusively measure agentic coding capabilities net of the confounders, we also ran Terminal-Bench with the latest 2.1 fixes available on GitHub, while increasing the timeout limits to 4 hours (roughly four times the 2.0 baseline). This brought the mean reward to 92.1%. (p188) > ...Mythos Preview represents only a modest accuracy improvement over our best Claude Opus 4.6 score (86.9% vs. 83.7%). However, the model achieves this score with a considerably smaller token footprint: the best Mythos Preview result uses 4.9× fewer tokens per task than Opus 4.6 (226k vs. 1.11M tokens per task). (p191)
- alyxya 6mo agoThe first point is along the lines of what I'd expect given that claude code is generally reliable at this point. A model's raw intelligence doesn't seem as important right now compared to being able to support arbitrary length context.
- zozbot234 6mo agoGood catch. If it's "too slow" even when ran in a state-of-the-art datacenter environment, this "Mythos" model is most closely comparable to the "Deep Research" modes for GPT and Gemini, which Claude formerly lacked any direct equivalent for.
- AlexC04 6mo agobut how does it perform on pelican riding a bicycle bench? why are they hiding the truth?! (edit: I hope this is an obvious joke. less facetiously these are pretty jaw dropping numbers)
- bertil 6mo agoWe are all fans for Simon’s work, and his test is, strangely enough, quite good.
- ninjagoo 6mo ago> Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) > Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% > GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% > MMMLU: 92.7% / 91.1% / — / 92.6–93.6% > USAMO: 97.6% / 42.3% / 95.2% / 74.4% > OSWorld: 79.6% / 72.7% / 75.0% / — Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen Opus 4.6 or GPT-5.4, I don't know what to make of the significant jumps on other benchmarks within these same categories. Training to the test? Better training? And the decision to withhold general release (of a 'preview' no less!) seems to be well, odd. And the decision to release a 'preview' version to specific companies? You know any production teams at these massive companies that would work with a 'preview' anything? R&D teams, sure, but production? Part of me wants to LoL. What are they trying to do? Induce FOMO and stop subscriber bleed-out stemming from the recent negative headlines around problems with using Claude?
- TacticalCoder 6mo ago> Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen We're not reading the same numbers I think. Compared to Opus 4.6, it's a big jump nearly in every single bench GP posted. They're "only" catching up to Google's Gemini on GPQA and MMMLU but they're still beating their own Opus 4.6 results on these two. This sounds like a much better model than Opus 4.6.
- ninjagoo 6mo ago> We're not reading the same numbers I think. We must not be. That's why I listed out the ones where it is barely competitive from @babelfish's table, which itself is extracted from Pg 186 & 187 of the System Card, which has the comparison with Opus 4.6, GPT 5.4 and Gemini 3.1 Pro. Sure, it may be better than Opus 4.6 on some of those, but barely achieves a small increase over GPT-5.4 on the ones I called out.
- nimchimpsky 6mo agobarely competitive ? Mythos column is the first column. You are the only person with this take on hackernews, everyone else "this is a massive a jump". Fwiwi, the data you list shows the biggest jump I remember for mythos
- johnnichev 6mo agodamn... ok that's impressive.
- WinstonSmith84 6mo agoNot discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly lower than GPT, and even lower than Gemini.
- muyuu 6mo agoWhen did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.
- gck1 6mo agoTried Gemini 2 weeks ago to see where it's at, with gemini-cli. Failed to use tools, failed to follow instructions, and then went into deranged loop mode. Essentially, it's where it was 1.5 years ago when I tried it the last time. It's honestly unbelievable how Google managed to fail so miserably at this.
- 4b11b4 6mo agoTheir harness might be behind
- gck1 6mo agoI think failures that I observed with gemini are unrelated to the harness. Because the same failures happened with third party harnesses too.
- unsupp0rted 6mo agoIt’s great on AI Studio. Harness issues, I agree.
- Kailhus 6mo agoI have not experienced any issues with Gemini 3.1 Pro.
- sandos 6mo ago
- matheusmoreira 6mo agoWow. Mythos must be insanely good considering how good a model Opus already is. I hope it's usable on a humble subscription...
- crimsoneer 6mo agoYou get a single call a month. Use it wisely.
- FridgeSeal 6mo agoWhat is the meaning of life, the universe, and everything? > Thought for 7.5 million years
- matheusmoreira 6mo agoHello, Claude! > Rate limit reached
- cesarvarela 6mo agoI thought they were bluffing when they talked about the scaling laws, but looking at the benchmark scores, they were not. I wonder if misalignment correlates with higher scores.
- maplethorpe 6mo agoFunny, I made my own model at home and got even higher scores than these. I'm a bit concerned about releasing it, though, so I'm just going to keep it local for now.
- casey2 6mo agoLots of benchmaxxing here. A few simple randomizations puts it back on par with gemini 3.1 and under 5.4 pro in most benchmarks
- mik09 6mo agoit might have broken a couple metrics since if you get above 90 percent it might be that the metric can not measure you well anymore right?