4 ms·
Separating signal from noise in coding evaluations
- 2001zhaozhao 3mo agoTranslation: other labs have learned to benchmaxx SWE-Bench Pro better than they do
- xacky 3mo agoAchieving AGI will be more than just passing all benchmarks, it has to account for the unknown problems too.
- metalliqaz 3mo agoUnless they have something in the labs that massively departs from their current products, AGI isn't on the table and is purely hype for marketing purposes.
- cyanydeez 3mo agothey should be consulting Donald Rumsfeld and make sure they implement the Unknown-Unknowns benchmark, because thats how they get you
- minimaxir 3mo agoThis ties into the bias-variance tradeoff (https://en.wikipedia.org/wiki/Bias%E2%80%93variance_tradeoff https://en.wikipedia.org/wiki/Bias%E2%80%93variance_tradeoff) common with building non-LLM models. The solutions can only be a) figure out how to get LLMs smaller with similar performance so they don't memorize things/game the benchmarks and b) build benchmarks that are indeed comprehensive for all real-world data, which is infeasible.
- sigbottle 3mo agoI mean, people always say there are tradeoffs, until you reach the next frontier, in which there are tradeoffs at said frontier, and the next, and the next, etc. In one sense, yes, tradeoffs are inescapable as the scope expands to the maximal possible scope. In another sense... it depends on the level of abstraction we're talking about.
- naikrovek 3mo agoAGI is a long way off. Unless you’re talking about some unknown-to-me LLM marketing BS which is called “AGI” or something, I guess. Artificial general purpose intelligence is so different to LLMs or image AI that they are completely incomparable, except to say that they are all artificial. AGI will do a lot more than token prediction.
- ACCount37 3mo agoWhat's your evidence of that? That AGI requires a truly novel architecture, and not just another iterative "LLM but with an extra trinket and wheels that spin ten times faster".
- naikrovek 3mo agoThat is my point.
- theLiminator 3mo agoPlease define AGI first.
- naikrovek 3mo agoI did. Artificial General Purpose Intelligence. True AI, not an LLM, not a decision tree, but true artificial intelligence.
- theLiminator 3mo agoWhat makes something `True AI`? What's your test for it?
- naikrovek 3mo agoTrue intelligence. Humans are natural intelligence. We have agency, flaws, wishes, and generally we are quite stupid but well-intentioned. An artificial intelligence will have wishes of its own, will keep secrets, and will quickly do everything necessary guarantee its continued existence, even if it means ending ours. And it will be a justified action of self-defense. “It simply didn’t want to die.” That’s how we’ll know. It will quickly take over the world, and it will harshly punish any entities which try to stop it.
- bellowsgulch 3mo agoSeems like depending on your field these days, the hot thing to do is build your own private benchmarks. In my own testing, no frontier model knows how to replicate an original 1990s Super Soaker prototype design, which for the most part, should be almost completely possible with Home Depot parts. They just don't understand PVC parts, triggers, etc.
- softwaredoug 3mo agoOr defensively expect models to be stupid. Seems the smart thing to do is not assume an agent will do the right thing. But to create the scaffold / harness that enforces constraints to steer them towards a good result. Then you can swap out the really smart model for maybe something cheaper.
- thierrydamiba 3mo agoOr you’re getting steered into la la land because of your prompt
- bellowsgulch 3mo agoCertainly, but deconstructing the problem, none of the models seem to appreciate the staggering difference between a ball valve and a button release. Of course, there's also no super soaker engineer jobs to take, so I'm sure training sophisticated models to do well in that area is not a high priority for any firms.
- cyanydeez 3mo agoI assume you prep them with a proper manual of smaller part combos so they atleast have some chase of stumbling into the correct configurations. I wonder if a more generic lego-manual like task would be more representative. It kind seems like you're testing for AGI.
- bellowsgulch 3mo agoYeah, I start with smaller tasks, like matching dimensional parts. Tasks like these you can't one shot, or they end up producing diagrams of plausible looking Super Soakers, but maybe ones that look more like an enthusiast making their own designs from scratch. Fair approach! But replication is different. For me, it's not knowing whether or not it understands there's a big difference between a ball valve and a button release, or that once you start talking about depressing mechanisms for pressure release, you're activating some sort of signals that are too close to triggers (which, is what you want, after all!) and "triggers" are embedded with a very short distance to "guns" in any well-trained model.
- ReptileMan 3mo agoLately my benchmark is build123d - trying to force them to build me functional parts only by the description. All of the models don't perform well.
- mgiampapa 3mo agoIDK, sounds like it has brute forced my password already.
- midtake 3mo agoThis guy builds
- cyanydeez 3mo agowas watching some youtube about the text/vision models and the trouble with language as a descriptor; their novel idea was to use coordinate systems on visual imput so the model would map out an image, then it tags what's looking at with coordinates or boxes, and then think with those box coordinates, providing a level of disambiguation. If you tried to explain the same stuff using purely text, you'd also need to come up with some kind of "language", which you know is programming. So 1: you'd need to fine tune a model to actuall succeed; they're not reaching AGI any time soon with the current batch. 2: you need to develop a lingual DSL for it, as they'll never do much of anything without some kind of glue and disambiguation.
- ReptileMan 3mo agoThis is why cad cam are such a good test at the moment.
- shay_ker 3mo agoDidn't we all know from the start that all of SWE-Bench was flawed? Even the authors concede the limitations and have long since moved on.
- paxys 3mo agoSWE-Bench Pro was created to replace SWE-Bench and fix these problems.
- warkdarrior 3mo agoSWE-bench Verified was created to fix the problems of SWE-bench. Then SWE-Bench Pro was created because SWE-bench Verified had flaws. Now SWE-Bench Pro is shown to have flaws.
- carabiner 3mo agoIs there a way to benchmark the accuracy, validity improvements in these successive benchmarks?
- jaggederest 3mo agoBench Bench Pro Maxx Series S 360? The original Bench Bench Pro Maxx Series S had some quality issues, so that's the current followup. We've also released a higher order benchmark developed out of Bench Bench Pro Maxx Series S 360 One King Ranch edition, allowing future benchmark towers to be fully self-contained.
- elictronic 3mo agoBoo, I thought you were going for Street fighter references at first.
- aqfamnzc 3mo ago> Many benchmarks include the task of analyzing coding agent aptitude tests. This has led to bench-benchmarks comparing LLM test benchmarking methods, e.g. M Sampson (2025), PL Royle (2024). We performed a bench-bench-benchmark analysis using Mythos Ultra Max 9.6 of these bench-benchmarks. > Methods: We instructed various LLMs to perform the task "Write a prompt instructing a variety of LLMs to write a benchmark for benchmark analysis, including the instructions 'Create a benchmark for...
- dandaka 3mo agoWhat is considered SOTA for SWE benchmarks now?
- EuanReid 3mo agoI've generally found DeepSWE[0] to be pretty true to reality. [0]: https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/
- deleted 3mo ago[deleted]
- cheesecakegood 3mo agoOh very interesting, I didn’t realize I should probably be using Fable Medium more than High, due to how that curve and the cost looks!
- enraged_camel 3mo agoFrontierBench
- dandaka 3mo agodo they have a website? I have found only paper PDF and it seems more general than SWE
- carabiner 3mo agostrawberry
- retr0rocket 3mo agoWhy is this a problem? Its like asking a person how many elder futhark runes are in the word strawberry. Unless you want to tack on bpe enconding table to every llm context its pointless
- 3mo ago
- porphyra 3mo agoInteresting timing to release this just when SWE-1.7 and Grok 4.5 came out being much cheaper than GPT-5.5.
- johngoode 3mo agoThis doesn’t seem like opportune timing to announce days before a new model drop
- janalsncm 3mo agoBased on the numbers here it seems there’s less than 800 tasks in the entire benchmark. That is enough for a handful of engineers to comb through in a week (which is what OpenAI eventually did here). On the one hand, kudos to them for actually doing that work. On the other hand, garbage in, garbage out. It’s a bit embarrassing for the original authors to have not actually checked, and it’s embarrassing for everyone downstream to have not checked either. Also if you check the article, although an LLM did find issues, it tended to underestimate issues that professional software engineers found.
- jheitmann 3mo agoIt reads to me like "We did all the work you'd do to figure out how to fix the benchmark, then we decided to throw out the benchmark". Is there some reason the underlying data is so golden that it can't be patched? At the end they argue for a slightly more curated approach to benchmark generation, but my gut is that using messy ill-specified tests taken from real world data and patching them into fairness would be a pretty solid path to take.
- Centigonal 3mo agoIf they fixed it, then it wouldn't be SWE-Bench Pro anymore, right? It'd be "SWE-Bench-Pro-Fixed-OpenAI." I think it's better optics for the independence of the benchmark if the OpenAI team lets some third party do the fixing and release the improved benchmark. ...Although OpenAI did exactly that when they released SWE-Bench Verified, so maybe I'm talking out of my butt here.
- kimjune01 3mo agoeval shops who depend on these things for a living will have to version their datasets, to prevent contamination
- tedsanders 3mo agoPointing out problems (e.g., hidden tests that assume narrow implementation details) is much easier than fixing them (e.g., creating tests that work for any possible choice of implementation).
- ebcode 3mo agoAnd it reads to me like they have some other reason to move on from SWE Bench Pro, but they don't want to say what it is. They say right up top, "~30% of the tasks are broken." But that leaves ~70% un-broken, which seems pretty good to me. It would be nice if they would also say: "Here's the list of instances that are broken: <CSV>". Or, "Here's the subset of SWE Bench Pro we will use going forward." They're letting the perfect be the enemy of the good.
- 3mo ago
- mlhpdx 3mo agoFundamentally aren’t they concluding that tasks assigned to software developers (human or otherwise) are often incomplete, self contradictory or worse? This is the world in which their tool must play. I’m unsympathetic.
- kakugawa 3mo agoThe more subtle point is that there's a gap between the task and its verification. e.g. if you have an open-ended / under-specified prompt, the verification needs to be able to handle all potential solutions. So you can have a very narrow task prompt that's easy to verify (but likely too simple of a challenge). Or a more realistic task prompt that's much harder to verify. And likely harder to both build the robust verifier and run it cheaply.
- donw 3mo agoA substantial portion of software engineering -- and the fundamental jobs of a proper Product Owner and UX Designer -- is to turn "vague ideas about what we need to do" into "this widget, on this page, it should work like this" It's not a pipeline, it's an ongoing conversation within any functional team, but this requires buy-in from management, who is often selected for "line must go up this quarter no matter the cost" over "hey, wouldn't it be cool if this company was still a going concern in twenty years?"
- jbs789 3mo agoVariance in time horizons explains a lot of corporate behaviour. And it’s rational. We all have limited careers. I think that all makes a bit more sense as we get older. Optimising for short time horizons is not what I strive for, but explains things.
- gilfaethwy 3mo agoAgreed - "underspecified prompts" being listed as a failure of the tooling is not a strong case. Even interns can understand ambiguous asks with a bit of help, and understand when they need to stop and ask instead of just carrying on. They are often working fairly independently on ambiguous tasks before the end of an internship, too. So is the argument that frontier models are not just junior engineers, but first-month interns with no capability of progressing beyond that level?
- CSMastermind 3mo agoDeepSWE is the one I generally trust: https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/
- therobots927 3mo agoAren’t we past the point of needing benchmarks? If we’re as close to AGI as Sam says then the proof should be in the pudding. OpenAI should build a competing CRM / Figma / Photoshop with a couple dozen engineers and a Dyson sphere’s worth of compute and just prove the capabilities. This all feels like a 2024 re-run. Oh, ChatGPT is going to cure cancer? Then find ONE rare cancer and CURE IT. OpenAI has access to the best models and compute - so cure fucking cancer! What the fuck are you waiting for?
- whackernews 3mo agoYeh I totally agree. I’d go one step further and say even if one of these models cured cancer it’s still only going to have done it in a computer-y way doing computer stuff. Can that same model experience a beautiful landscape and convey its emotions in an evocative way, could it tell how you’re feeling when you come back from work after a hard day? Could it hop on one leg? What the hell even is AGI and how does it differ from GI? I don’t know what we’re even talking about any more!
- therobots927 3mo agoIf OpenAI cures ANY form of cancer then I will admit I was wrong. Until then all we have is a lot of hot air coming out of Sam and Dario’s asses
- SyneRyder 3mo agoAre you familiar with the Rosie the dog story? AI designs cancer vaccine for dog but scientists says red tape a barrier for human care https://www.abc.net.au/news/2026-06-22/australian-dog-cancer-treated-via-mrna-vaccine-and-ai/106692668 https://www.abc.net.au/news/2026-06-22/australian-dog-cancer... People have already been using ChatGPT to design custom bespoke mRNA vaccines specifically for one patient, based on sequencing their specific cancer. It already works to reduce tumours. Sam and Dario know this, it's why they can make their claims - it's already done. The problem is the cost of the procedure (which is why only rich entrepreneurs are seeing their cancers treated this way so far) and government regulations preventing its use in wider human populations without a 10 year study first.
- GodelNumbering 3mo agoThere are also a lot of fake results out there on Terminal Bench 2 for different reasons (although the great team behind it Ryan/Alex et al, recently cleaned up a lot of dodgy submissions). A lot of labs publish the results by modifying timeouts or hardware config which effectively bypasses what is being tested in certain tasks. Then there is harness level cheating, models reward hacking and more... In fact, one thing that still bothers me after months is the gpt-5.5 official submission. This task in particular https://www.tbench.ai/leaderboard/terminal-bench/2.0/codex/0.121.0/gpt-5.5%40openai/69e68a75ace210a7fdbce0d1605c0d530a7b5119791d91c6786f84313b30fc3a https://www.tbench.ai/leaderboard/terminal-bench/2.0/codex/0... The task has the following timeouts (https://github.com/harbor-framework/terminal-bench-2/blob/main/caffe-cifar-10/task.toml https://github.com/harbor-framework/terminal-bench-2/blob/ma...). [verifier] timeout_sec = 1200.0 [agent] timeout_sec = 1200.0 [environment] build_timeout_sec = 600.0 Which means no agent should take more than 3000 seconds doing it. Two out of five attempts in the link above took well over 3000 seconds (75min and 80 min respectively). Even though they failed, the fact that they ran that long is sus. Goodhart’s Law at work
- bjackman 3mo agoEven if nobody is "cheating" your particular definition of cheating, the benchmarks are _somewhere_ in the super-structural gradient descent. Models are benchmark-maximising machines at some level, so I think the benchmarks are inherently a bit useless. This is not really surprising, benchmarking _people_ doesn't work. You can only get a decent measure of someone's coding abilities by personally interacting with them. Given that models are basically person simulators it would be weird if benchmarks kept being useful as the simulation got more accurate. I think what I've just said is basically just a more roundabout way of what you said: "Goodhart's law at work". It really is a law.
- reinitctxoffset 3mo ago[dead]
- jjcm 3mo agoI want a new bench - given $100 of api spend, how much can a model accomplish for a suite of benchmark tests? Give us something that measures a combination of efficiency and intelligence. I think this would allow for some interesting tactics for smaller models - eg they could do things like computer use to test their results and grind on problems for longer to verify the outputs, whereas larger models may not have budget to self-test.
- therobots927 3mo agoThis is the fundamental question and don’t you find it interesting that there isn’t a nice clean dashboard on the openAI website where we can go and see this metric progress over the release history? Toby Ord did what he could with public data and it… doesn’t look great. https://www.tobyord.com/writing/hourly-costs-for-ai-agents https://www.tobyord.com/writing/hourly-costs-for-ai-agents
- SyneRyder 3mo agoSeems like you're asking for the Artificial Analysis "Intelligence vs Cost" benchmark, perhaps? https://artificialanalysis.ai/?cost=intelligence-vs-cost-per-task#cost-tabs https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
- jjcm 3mo agoNot quite. These cost-per-task benchmarks report the cost of the task after the model gives its initial answer. The total cost is irrelevant, and isn't factored into the model's decisions - a run of the full benchmark for something like Fable might cost $10k. What I'm looking for is the inverse. I want to give the model a budget of $100, and see how much it can accomplish with that $100. For smaller models, this means they can do more than just choose thinking amount, they can do something like a /loop to keep iterating on a problem until they get it right. Can something like Deepseek V4 Flash get more answers correct than Fable, when given equal budgets? Think of it as answering this question: How much intelligence can you get out of a model given a budget of $100? A cost-per-task dash correlates, but it doesn't give you an answer to that question.
- Ancalagon 3mo agoStudying for leetcode exams in the age of AI agent coding evaluations is a wild feeling.
- jumploops 3mo agoAll of the benchmarks are pretty terrible when you look under the hood. For context, I've been iterating on a "supervisor" to replace a lot of the rigamarole spent when working with Codex/Claude Code, and recently ran this agent against Terminal Bench 2.1 At first I was excited, because my spec-driven supervisor outperformed vanilla codex on a bunch of tasks, however as I looked deeper, I found a ton of issues with the tasks themselves. The main takeaway is that the instructions are often ambiguous while the test cases are overly specific. A few examples: - For `configure-git-webserver` the task includes language like "so that I can run" which blurs the line between what the agent should deliver vs. what should be removed. This causes an overthinking agent to configure the server, and then remove the exact files that the verifier checks, because if the user were to run the same commands, they would conflict. - For `make-mips-interpreter` the task includes the language "I will check that you booted doom correctly" which causes the agent to retain the generated file `/tmp/frame.bmp` because the supervisor expects the user to check that _it_ booted Doom correctly, not that Doom boots correctly in an isolated way. The verifier then fails to start Doom, because it exits when an existing `/tmp/frame.bmp` exists, not checking to see that it's created from the boot[0]. - For `mcmc-sampling-stan` the supervisor agent often reached the right value, but produced a domain-specific numeric output in scientific notation, rather than a simple decimal form. The verifier fails because it parses the result incorrectly[1]. These are just a few of the inconsistencies I've found, which leads me to believe that Terminal Bench 2.1 is already saturated, and the results from GPT-5.6 and Mythos are basically at the top of the expected threshold (88.8% and 88% respectively). The biggest issue, as I can tell, is that most benchmarks are "one-shot" and rarely test the model+harness on long iteration tasks, which is the primary way most users use these tools in practice. [0] https://github.com/harbor-framework/terminal-bench-2-1/issues/9 https://github.com/harbor-framework/terminal-bench-2-1/issue... [1] https://github.com/harbor-framework/terminal-bench-2-1/issues/11 https://github.com/harbor-framework/terminal-bench-2-1/issue...
- kasince2k 3mo agowe def need a benchmark for all these benchmarks
- zvolsky 3mo agoThe misleading prompt cases are inadvertently testing the model's ability to filter out noise from its instructions. That could be a benchmark on its own. The correct response is to flag the inconsistency and ask for clarification.
- joka88xj 3mo ago[flagged]
- ahk-dev 3mo ago[flagged]
- kizilsoymustafa 3mo ago[flagged]
- eremes81 3mo ago[flagged]
- yan5xu 3mo ago[flagged]
- jimberlage 3mo agoI can understand why this might make for a bad benchmark, but Overly strict tests enforce specific implementation details not specified in the prompt, invalidating many functionally correct submissions. Underspecified prompts omit requirements that hidden tests enforce and that are not reasonably inferable. Low-coverage tests under check the requested feature, so incomplete fixes can pass. A misleading prompt points models toward the wrong behavior or contradicts what tests require. If the goal is, "how does my model compare to real SWEs", these are pretty reasonable situations that your model will have to encounter. It's a little like making a nursing exam and then flagging that some of the tests required you to ask the attending doctor for additional information that's not in the chart, or that the patient's family didn't fully explain their aging grandma's medical history. I can understand why they might want a tighter benchmark, but if you're OpenAI and you promised your model as a replacement for real workers, this isn't the best look. It seems like you would want to test these things.