8 ms·
SWE-bench Verified no longer measures frontier coding capabilities
- enesz 5mo ago[dead]
- w4yai 5mo agoI don't understand these websites which force translation to my native language. I mean, it's fine as it's useful for many people, but where is the button for disabling it ? Or why is it enabled by default ? "codage de pointe" sounds so weird and cringe in French.
- Toutouxc 5mo agoSame for apps and games. I understand English just fine, no need to switch to your shitty Google-translate localization just because my iPhone or PlayStation is set to my native language.
- LukaD 5mo agoDoes your browser request French via an Accept-Language header perhaps? What really infuriates me is when sites don’t respect that header and give you a translation based on IP location.
- embedding-shape 5mo agoRegardless if it does or not, users should be able to manually override what language the website is in, at least be able to read the native one, regardless of what the original language was, what headers you send and where geodatabases think your IP is from.
- w4yai 5mo agoCorrect answer! What a bad UX
- deleted 5mo ago[deleted]
- 1a527dd5 5mo agoThis feels very much like "we are now moving the goal posts".
- neversupervised 5mo agoBut this is the good kind of goalpost moving
- iLoveOncall 5mo agoOnly if you didn't read the article. They're saying they need to move on from it because the benchmark is flawed (without bringing in proof) and that's why they can't hit 100%. It's not a "our models are so good that the benchmark is too easy" thing.
- f33d5173 5mo ago> without bringing in proof Did we read the same article?
- embedding-shape 5mo agoI feel like they're quite open about why they think the benchmark doesn't work anymore: > We also found evidence that models that have seen the problems during training are more likely to succeed, because they have additional information needed to pass the underspecified tests. > This means that improvements on SWE-bench Verified no longer reflect meaningful improvements in models’ real-world software development abilities. Instead, they increasingly reflect how much the model was exposed to the benchmark at training time.
- MattRix 5mo agoHow can you say “without bringing in proof” when there is literally proof in the article?
- MattRix 5mo agoOnly if you didn’t read the article…
- neversupervised 5mo agoTerminal Bench is the future
- embedding-shape 5mo agoFirst, you might want to say why you think so, otherwise this is just borderline spam. Secondly, when your praise things (without motivation or reasoning even), and you've contributed to that specific thing, please say that up front instead of just praising the thing, again it makes it look like spam otherwise.
- vintagedave 5mo ago> We audited a 27.6% subset of the dataset that models often failed to solve and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions, despite our best efforts in improving on this in the initial creation of SWE-bench Verified. Is this saying a quarter* of the questions and answers were wrong, this whole time?! If so, how was this ever, in any way, a valid measurement? And what was the process for creating this benchmark and how did it end up with such an extraordinarily poor set of data? (There is a description later of how, which seems to be a high standard and I struggle to understand how it aligns with the other results they discuss.) Kudos to them for highlighting the issues, but I am left with questions. [*] Not one in four, but one in six, thanks commenters for the correction; leaving the original since, eh, my bad, and it lets replies make sense. I feel the broad point still stands!
- motoboi 5mo agoIt’s saying that 16% of the problems have well, problems.
- vintagedave 5mo agoYou're right - I did not apply the math. (I won't edit, in order to let the parent comment still make sense, and thankyou for the correction.) So not one in four, but one in six problems have problems. That is extraordinarily high and the point still stands: is this truly saying a [large proportion] of the questions and answers were wrong, this whole time, and if so how was it ever a valid measurement?
- motoboi 5mo agoWait until you discover how many wrong labeled images in imagenet and that it still kickstarted the deeplearning revolution.
- deleted 5mo ago[deleted]
- adityamwagh 5mo ago> We also found evidence that models that have seen the problems during training are more likely to succeed, because they have additional information needed to pass the underspecified tests. No shit, Sherlock!
- djoldman 5mo ago> We have incorporated these findings into our recent evaluation efforts. In the last months we’ve chosen to report results from the public split of SWE-Bench Pro. We recommend other model developers do the same. SWE-bench Pro is not perfect, but empirically seems to suffer less from contamination issues. https://arxiv.org/pdf/2509.16941 https://arxiv.org/pdf/2509.16941
- Jcampuzano2 5mo agoIts pretty clear that any benchmark that comes out will be outdated and exist within the training data with short measure. There will always be an incentive to optimize specifically for these benchmarks even if just for marketing material. Sure there is a training cutoff, but its usually only 3-6 months off of the public release dates. The problem with coding benchmarks then becomes creating novel benchmarks that are guaranteed to not already be in the training data, and not borrow anything from previous benchmarks. In this regard I don't think any benchmark that was created before a given model is released should ever be considered valid or representative of model performance. The potential financial gain for including the data just to be able to market a minor improvement is too swaying. With that in mind they should honestly just stop including benchmarks altogether in marketing material Let the model speak for itself and let the community decide, but of course that will never slide with corporate types with so much money on the line.
- mnky9800n 5mo agoThis is why I made Zork bench. Zork, the text adventure game, is in the training data for LLMs. It’s also deterministic. Therefore it should be easy for an LLM to play and complete. Yet they don’t. Understanding why is the goal of Zork bench. https://github.com/mnky9800n/zork-bench https://github.com/mnky9800n/zork-bench
- kqr 5mo agoI have worked on similar problems. See e.g. [1]. The LLMs I have tested have terrible world models and intuitions for how actions change the environment. They're also not great at discerning and pursuing the right goals. They're like an infinitely patient five-year old with amazing vocabulary. [1]: https://entropicthoughts.com/updated-llm-benchmark https://entropicthoughts.com/updated-llm-benchmark (more descriptions available in earlier evaluations referenced from there)
- mnky9800n 5mo agowe should talk. i sent you an email.
- Jimmc414 5mo agoGoodhart’s Law in reverse, what can’t be gamed gets rejected.
- cbg0 5mo agoSWE-bench verified was created in collaboration with OpenAI. It's also an open dataset so prone to contamination, meaning it can be gamed.
- stephen_cagle 5mo agoYou've almost buffer overrun Goodhart's Law into the https://en.wikipedia.org/wiki/McNamara_fallacy https://en.wikipedia.org/wiki/McNamara_fallacy . :]
- varispeed 5mo agoIssue with these benchmark also is that they measure a model you are unlikely going to be routed to. My experience with Anthropic is that despite using Opus 4.6 and 4.7, most of the time the performance is matching low B parameter Qwen. I think there should be a way to verify what model is actually being used to process prompts - that should be independently verified. At the moment it is so bad, you have to ask verification question to the model in form of a non-trivial problem. If it solves it, then there is a chance you actually get Opus and not an impostor and so you can continue the session instead of restarting it hoping you get routed correctly. But that does not help if model is replaced with cheaper one mid session. I've got so much work lost because of these shenanigans.
- alansaber 5mo agoI'm sure some inference providers don't, but most intentionally obfuscate this data. They have the full trace logs- my impression is that they don't share them because it's their competitive advantage, and it's easier for a competitor to distil their model if they did.
- gruez 5mo ago> My experience with Anthropic is that despite using Opus 4.6 and 4.7, most of the time the performance is matching low B parameter Qwen. Is this just the next level of the "they're serving quantized models!" theory?
- varispeed 5mo agoNot a theory buy lived experience. You never know when you get the nerfed session.
- gpm 5mo agoCuriously Opus 4.7 claims to have a 87.6% pass rate and Mythos claims to have a 93.9% pass rate... leading to the conclusion that it's actually possible to "solve" the problems that OpenAI claims are incorrect.
- 2ndorderthought 5mo agoOr that opus and mythos are training on the data somehow such that there solutions are incorrectly right. Or that openai is lying/wrong. Or that all of these companies are cheating so much it doesn't really matter and never did.
- MattRix 5mo agoThe problem isn’t that the tasks are impossible to solve, it’s that they’re underspecified and/or impossible to solve consistently (ex. because a test is expecting the solution function to have a specific name that wasn’t specified in the task itself). So maybe Anthropic runs Mythos through the benchmark 10000 times and takes the highest score, who knows?
- gpm 5mo agoWe actually know that a "100% pass rate" is trivially possible: https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/ https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/ Anthropic p-hacking the benchmark strikes me as cheating, and somewhat unlikely. Mythos figuring out how to cheat at the benchmark strikes me as much more likely. But if that hypothesis is the explanation the interesting part is Opus 4.7 (but not 4.6) seems to be doing the same.
- gruez 5mo ago>Mythos figuring out how to cheat at the benchmark strikes me as much more likely. Define "cheat". If it's just hacking the test harness to return "PASSED", surely this would be easily detected with some human auditing? It sounds far more likely their solution are designed to pass the incorrect tests. That might be considered bad in a SWE context, but it's not exactly cheating either. It might even be considered a good thing, eg. in the context of backwards compatibility. [1] https://learn.microsoft.com/en-us/troubleshoot/microsoft-365-apps/excel/wrongly-assumes-1900-is-leap-year https://learn.microsoft.com/en-us/troubleshoot/microsoft-365...
- ripvanwinkle 5mo ago>>In our analysis we found that all frontier models we tested were able to reproduce the original, human-written bug fix used as the ground-truth reference, known as the gold patch, or verbatim problem statement specifics for certain tasks, indicating that all of them have seen at least some of the problems and solutions during training this statement alone seems to invalidate the SWE-bench tests
- threepts 5mo agoWhy don't they ask their premier model to generate a bench for them? Jokes aside, a benchmark I look forward to is ARC-AGI-3. I tried out their human simulation, and it feels very reasoning heavy. Leaderboard: https://arcprize.org/leaderboard https://arcprize.org/leaderboard (Most premier models don't even pass 5 percent.)
- alansaber 5mo agoVery (reasoning) heavy benchmarks do seem like the way to go, being the hardest to game.
- falcor84 5mo agoThey focus on minimizing the number of moves and don't allow any harness whatsoever, putting the bar extremely high. The current top verified contender (Claude Opus 4.6) is at only 0.45%. But with how new it is, I expect a lot of improvement in the next generation of models.
- threepts 5mo agoOptimal for judging actual reasoning ability rather than an LLM's ability to regurgitate knowledge from a necropost on HN/Reddit/Twitter from 2018.
- knollimar 5mo agoa small harness that stores text files and manages context could be useful, otherwise you lose all ability to measure that skill (and that's important because it represents real world use cases on large code bases)
- anthonypasq 5mo agoarc agi isnt testing a models ability to store files and code things. its testings its ability to reason through puzzles given the same information as a human
- retinaros 5mo agoit never did
- gertlabs 5mo agoA better benchmark needs to be objectively scored, have multi-disciplinary, breadth, and be scalable (no single correct answer). That's what we designed at https://gertlabs.com https://gertlabs.com. We put a lot of thought into it, and kept it mostly (not fully) related to problem solving through coding.
- orangebread 5mo agoWow. This benchmark definitely feels more accurate than the other rankings I've seen. My experience with gpt 5.4/5.5 is that they are technically flawless and if there are any technical issues that is because the input didn't provide enough clarity; that's not to say that it doesn't autonomously react to any issues during bug fixes or implementations, but it'll tend to nail its tasks without leaving behind gaps. Opus otoh is overrated in terms of its technical ability. It is certainly a better designer/developer for beautiful user experiences, but I'll always lean on gpt 5.5 to check its work. The biggest surprise in the benchmark is Xiao-Mi. I haven't tried it yet, but I will be after looking at this. Grats on your team for putting together something meaningful to make sense of the ongoing AI speedrun! Great work!
- gertlabs 5mo agoMuch appreciated! MiMo V2.5 Pro is by far the most underrated recent release (probably because it wasn't open weights from the start).
- euleriancon 5mo agoAre we looking at the same data? On that site I see that opus 4.7's and gpt 5.5's g scores are within each others confidence intervals, and both significantly ahead of the number 3 model. Your comment makes it sound like they are miles apart, which the benchmark doesn't seem to support. Edit: I looked at the data more and the two models are only basically equal when looking at the mean of all the tests. Gpt 5.5 significantly outperforms opus 4.7 in coding, while opus 4.7 significantly outperforms in "decision making." I'm not seeing details on what decision making explicitly means.
- DeathArrow 5mo agoSo we need to generate benchmarks after the models finish training. Or we need to keep the solutions to the benchmark problems as closed source.
- kqr 5mo agoIt was never that great, it seems. For all of 2025 there was virtually no improvement in the rate at which models produced quality code. They only got better at passing automated tests. https://entropicthoughts.com/no-swe-bench-improvement https://entropicthoughts.com/no-swe-bench-improvement
- civvv 5mo agoThis is likely true. I think model quality has stagnated and that its likely a non-trivial task to find a new improvement vector. Scaling the width of the model (which has been the driving force behind the speed of improvement thus far) seems to have reached its limit. It will be interesting to see the implications of this. Tooling can only do so much in the long term.
- mxwsn 5mo agoHow do you know that width scaling has been the driving force of improvement?
- waterTanuki 5mo agoI mean, it's not exactly a PhD level question. One can infer from the extreme demand of GPUs and DRAM + new data center construction that all the providers are banking on width.
- svnt 5mo agoNo? That could just be fomo, actual adoption, or a number of other things.
- civvv 5mo agoI am no insider and have never even tried to build an LLM, so I can only guess. But the general sentiment seems to be that this is the case. If you are interested, I would recommend you read the MIT paper "Superposition Yields Robust Neural Scaling" [0]. It confirms an interesting trend: models represent more features/concepts than they have clean independent dimensions, so features overlap. Increasing model dimension reduces this geometric interference, which lowers loss in a predictable way, but with diminishing returns. This has, in my opinion, likely been the primary vector in getting better models thus far, but MIT mathematically proves that it yields diminishing returns for each new dimension added. It will get more and more expensive and the cost-return will or probably already has made it infeasible. Ilya appear to support sentiment this as well. [1] [0] - https://openreview.net/forum?id=knPz7gtjPW https://openreview.net/forum?id=knPz7gtjPW [1] - https://www.businessinsider.com/openai-cofounder-ilya-sutskever-scaling-ai-age-of-research-dwarkesh-2025-11 https://www.businessinsider.com/openai-cofounder-ilya-sutske...
- techpulselab 5mo ago[dead]
- DeathArrow 5mo agoSo Opus 4.7 and Mythos are solving problems that are impossible to solve?
- tedsanders 5mo agoWhether a problem is "good" or "bad" is not always objective or simple. For example, you can have problems that are underspecified, with hardcoded tests for a particular solution (out of multiple possible solutions). If your solution works fine but used a different function name than the one hardcoded in the tests, you can unfairly score 0. When an eval has underspecified problems like these, you can still score 100% if you remember the original solution from your training data or if you just have taste similar to the original human authors. And both of these qualities - good memory and good taste - are great, but they'll be rewarded unfairly relative to a model that still did exactly what it was asked but in a different way than the hardcoded tests expected.
- karmasimida 5mo agoTo some extent yes. It is not impossible to solve in absolute terms, in the sense, all necessary pieces of information are presented in the repo + problem statement. But it is impossible to solve in the sense, unless you read the ground truth, you are NOT able to solve it the way the test patch demands. Simply not plausible to me that model can read the problem statement so precisely that it nails exactly, like 100% what the test suite is trying to test.
- rustyhancock 5mo agoI think an Olympiad format is better. But the financial incentive is such that it might be near impossible to stop leaks. I.e. A panel comes up with a series of problems. Like advent of code or project Euler but more complex and constricted. Benchmark outcomes could be performance points and measure of cost, time to solution (well token count really). A couple times per year it's run. It avoids overfitting. Overtime the tasks can become more complex if needed. If they benchmax it into being able to complete full products from spec and robust implementations amazing.
- cjsaltlake 5mo agoSWE-bench was created to replace olympiad coding benchmarks. I think past olympiad coding benchmarks were much worse representative of real-world coding than something like SWE-bench, which is derived from real units of labor. Further, olympiad style benchmarks are arguably easier to contaminate / memorize unless you refresh it regularly; but that goes for SWE-bench too.
- rustyhancock 5mo agoI was picturing one-shot performance only for the benchmark, on novel real world tasks. I.e. the score on the March Olympiad you got in April isn't relevant. Simple enough that anyone could run it with a regular subscription. Really unless we can get the providers to ditch the gameable benchmarks they won't. But industries love nothing more than a benchmark they can manipulate.
- cowartc 5mo agoThe headline leads with contamination, but buried is that 59% of audited failures had test design defects. That's a measurement system never validated against ground truth before being adopted industry-wide as a score that mattered. They reported on it for two years but the gauge was broken the entire time.
- nothinkjustai 5mo agoAi comments are banned here.
- neuroelectron 5mo agoIt's really naïve to think any of the big AI companies won't cheat
- ryguz 5mo ago[dead]
- huflungdung 5mo ago[dead]
- languid-photic 5mo agoIt’s very hard to encode the properties that matter most in code in tests. [1] [1] https://voratiq.com/blog/your-workflow-is-the-eval https://voratiq.com/blog/your-workflow-is-the-eval
- parentheses 5mo agoThe timing makes me wonder if this is a direct response to Deepseek V4 having performance comparable to SOTA models.
- osti 5mo agoThis was published two months ago. Even though it was at a time that open source models are publishing comparable swe bench scores.
- cpard 5mo agoBenchmarks/evals are really hard and they become harder when there’s huge incentive to game them at an industry scale. ELT-Bench is another recent example. It was the first serious attempt at a benchmark for data engineering workloads, published about a year ago. A few days ago, a follow-up paper from a group that includes one of the original authors audited the benchmark itself. The team gfound that the benchmark has structural issues that biased results. Here’s the paper: https://arxiv.org/abs/2603.29399 https://arxiv.org/abs/2603.29399 None of these are new though, the industry has gone through all that before just in a smaller scale and there’s a lot to learn from that. Here’s a post I wrote on the parallels we see today to what happened with the benchmarketing wars of the database systems. https://www.typedef.ai/blog/from-benchmarketing-to-benchmaxxing-what-40-years-of-database-evals-can-teach-data-leaders-about https://www.typedef.ai/blog/from-benchmarketing-to-benchmaxx...
- fnordpiglet 5mo agoDatabase benchmarks are another. I have empirical experience though building classifiers that can have no precision measurement because the classifier performs invariably better than humans. They become the state of the art benchmark themselves and can’t be benchmarked except against themselves. These are for tasks that are non trivial and complex, but less logical than coding and less sustained reasoning. There may come a day though, when there is no calibrated benchmark that is independent of the models it’s measuring.
- softwaredoug 5mo agoIt’s just hard to make them not part of the training data. We see this a bit with BrowseComp plus and other deep research datasets. Not because frontier labs are trying to cheat, but just from training on the full web. You need new datasets perpetually.
- stavros 5mo agoOr hidden benchmarks, though it's then harder to get people to trust the results.
- wredcoll 5mo agoThis is somewhat tangential, but I want a model that can detect physical objects placed on top of a board from a picture/video, specifically warhammer 40k models. I want a model that can detect the actual units/models that are placed on top of the terrain/board so I can track how the models move during the game, but trying gemini and chatgpt they were absolutely rubbish.
- z33k 5mo agoAmiibo and Skylanders detect the pieces with NFC. Wiring up the whole board/ terrain with NFC readers would probably be difficult, though.
- addaon 5mo agoThe other classic approach has been a single camera under the table, but that conflicts with terrain use. mmWave radar is probably good enough for to localization at this point, and cheap, but distinguishing pieces is hard.
- wredcoll 5mo agoAn interesting thought but at the moment I was just talking about analyzing a video lol
- alphainfo 5mo ago[dead]
- vdalhambra 5mo ago[dead]
- lmeyerov 5mo agoIt's been fun benchmarking AI investigations at botsbench.com . Part of it is checking for these kinds of issues - we recently started seeing contamination in our first generation challenge, and less obvious, agent sandbox escapes for other kinds of cheating. Fun times!
- swyx 5mo agomore context in small writeup + we interviewd the team behind this when it was announced: https://www.latent.space/p/swe-bench-dead https://www.latent.space/p/swe-bench-dead
- ofirpress 5mo agoI'm a co-creator of SWE-bench: 1. SWE-bench Verified is now saturated at 93.9% (congrats Anthropic), but anyone who hasn't reached that number yet still has more room for growth. 2. SWE-bench Multilingual and SWE-bench Multimodal (which we'll open source in the next month) are still unsatured. 3. All benchmarks and benchmark paradigms eventually become saturated. That's why the SWE-bench team has worked hard on building the next stage of benchmarks, and we have a few that are already out, for example https://codeclash.ai/ https://codeclash.ai/ or https://algotune.io/ https://algotune.io/ . And we'll have more to say soon :)
- energy123 5mo ago> 93.6% (congrats Anthropic) But the article says "We audited a 27.6% subset of the dataset that models often failed to solve [which is 19.1% of the problems at time of publication] and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submission" 0.191 * 0.594 > 1 - 0.936 Does this mean that the audited subset wasn't representative? Or that Anthropic is getting high answers through some shady means?
- cjsaltlake 5mo agoI suggest reading the Mythos report's discussion on SWE-bench and contamination. I think it's fairly convincing that you can account for contamination and still trust SWE-bench numbers on models that aren't over-optimized for it.
- eugenekolo 5mo agoWithout SWE-Bench though, how will AI models properly game their results to show ~5-10% gain each iteration? Once a benchmark is known and there's billion of dollars on the line, obviously every company will game them.
- tripleee 5mo ago[dead]
- marlburrow 5mo agoThe "private benchmarks" suggestion comes up every time, but I think there's a more interesting axis: benchmarks built on top of already-public, already-stable test instruments. SWE-bench is fundamentally a corpus that lives on GitHub — once it ships, it leaks into training data automatically. Benchmarks built on contested qualitative instruments (psych tests, opinion surveys) have a different contamination profile because the correct answer doesn't exist in the training corpus to memorize — only the question does. That doesn't help for measuring coding ability specifically (you fundamentally need a code-correctness oracle), but for capability axes where the "answer" is a stated position rather than a verifiable fact, public + stable can still be useful. The SWE-bench problem isn't really "public", it's "public + has a fixed correct answer".
- axpy906 5mo agoOnce the bench is public it’s out and probably in the training data. Better to have your own and test it on a new model.
- hibouaile 5mo ago[dead]
- chhxdjsj 5mo ago[dead]
- jddj 5mo agoFor the most part I think we get the benchmarks we deserve. Many SWE-bench passing PRs would not be merged: https://news.ycombinator.com/item?id=47341645 https://news.ycombinator.com/item?id=47341645 Top model SWE bench scores may be skewed by git history leaks: https://news.ycombinator.com/item?id=45214670 https://news.ycombinator.com/item?id=45214670
- zachdotai 5mo agoI wrote about this recently here: https://fabraix.com/blog/adversarial-cost-to-exploit https://fabraix.com/blog/adversarial-cost-to-exploit I think the core issue is in static benchmarks and the community needs to start moving beyond measuring pass/fail (which worked when agents were incapable of doing much of the work) to dynamic evals that simulate more how we evaluate humans.
- kimjune01 5mo agoAI labs should compete on a bench that's adversarial, such as go or Starcraft
- pylonpeng 5mo ago[flagged]
- gmerc 5mo agoTranslation: Now that all rest sets are ingested, we need to move the bar that gave use several years of free PR. See also: https://this.os.isfine.org/blog/posts/us-ai-labs-love-the-ai-race-so-much-they-d-like-the-government-to-kneecap-their-/ https://this.os.isfine.org/blog/posts/us-ai-labs-love-the-ai...
- tokenhub_dev 5mo ago[dead]
- pkoiralap 5mo agoThis was bound to happen either organically or inorganically. Make sure it performs well on the benchmarks. And it doesn't really matter if it doesn't generalize outside of it right? :D Also similar: Graduate student descent. https://sciencedryad.wordpress.com/2014/01/25/grad-student-descent/ https://sciencedryad.wordpress.com/2014/01/25/grad-student-d...
- flowdesktech 5mo ago[dead]
- getverdict 5mo ago[dead]
- deleted 5mo ago[deleted]
- conorliu 5mo ago[dead]
- jimmypk 5mo ago[flagged]
- mohamedabdallah 5mo ago[dead]
- atlasagentsuite 5mo ago[dead]