10 ms·
Measuring AI Ability to Complete Long Tasks
- nrhrjrjrjtntbt 10mo agoWhy measure in minutes and not tokens? Seems you could cheat by slowing the ai down.
- wmf 10mo agoThey measure the time it takes a human to complete the task. They don't care how long the AI takes (although in practice it's much faster than human). Measuring tokens isn't a good idea because newer models can complete tasks using fewer tokens.
- grim_io 10mo agoThis seems like a good way to measure LLM improvement. It matches the my personal feeling when using progressively better models over time.
- Dwedit 10mo agoOpus is already the name of an audio codec.
- GaggiX 10mo agoOpus: "an artistic work, especially one on a large scale." The names Haiku, Sonnet, and Opus have not been chosen randomly.
- oidar 10mo agoAnd so much more intuitive than the OpenAI names for their models. I still don't get their naming scheme.
- p1esk 10mo agoHave you been living under a rock?
- pants2 10mo agoGemini is already the name of a Greek god, a constellation, a space mission, a crypto exchange, an astrological sign, a car, and a comic villain! How will we ever figure out which one someone is talking about?
- eleventen 10mo agoI recently asked Opus to just “Add vector search” to my current hobby project, a topic I know very little about. It set up manticore, pulled an embedding model, wrote a migration tool for my old keyword indices, and built the front end. I’m not exaggerating much either: the prompt was the length of a tweet. I think it would easily have taken me 4+ hours to do that. It ran in 15 minutes while I played Kirby Air Riders and worked on the first try. Afterward, I sort of had to reflect on the fact that I learned essentially nothing about building vector search. I wanted the feature more than I wanted to know how to build the feature. It kept me learning the thing I cared about rather than doing a side quest.
- ModernMech 10mo agoThe result of you having worked 4 hours to implement the thing is not just that you have the thing, it's that you have the thing and you understand the thing. Having the thing is next to useless if you don't understand it. At best it plods along as you keep badgering Claude to fix it, until inevitably Claude reaches a point where it can't help. At which time you'll be forced to spend at least the 4 hours you would have originally spent trying to understand it so you can fix it yourself. At worst the thing will actively break other things you do understand in ways you don't understand, and you'll have to spend at least 4 hours cleaning up the mess. Either way it's not clear you've saved any time at all.
- OxfordOutlander 10mo ago> inevitably Claude reaches a point where it can't help. Perhaps not. If LLMs keep getting better, more competent models can help him stay on top of it lol.
- evklein 10mo agoYou're still captive to a product. Which means that when CloudCo. increases their monthly GenAI price from $50/mo. to $500/mo., you're losing your service or you're paying. By participating in the build process you're giving yourself a fighting chance.
- yismail 10mo agoWould be interesting to see Gemini 3.0 Pro benchmarked as well.
- PunchTornado 10mo agoExactly. I don't understand how an article like this ignores the best models out there.
- cubefox 10mo agoThis article was published a long time ago, in March.
- yismail 10mo agoThat's true, but it looks like it's been updated since then because the benchmarks include Claude Opus 4.5
- alexgotoi 10mo ago[dead]
- simonw 10mo agoI didn't really understand the "long task" thing until I actually experienced it. The problem is finding a task you can set an agent that justifies working for that long. I finally hit one when I tried porting that Python HTML5 parser to JavaScript by pointing Codex CLI at the 9,200 html5lib-tests test suite: https://simonwillison.net/2025/Dec/15/porting-justhtml/ https://simonwillison.net/2025/Dec/15/porting-justhtml/ It's pretty amazing to watch tools-in-a-loop crunch away for >4 hours to solve a generally difficult problem through sheer brute-force.
- ehnto 10mo agoI think you might be misunderstanding the article actually, this is about AI solving tasks as measured by how long it takes a human to solve the task. The AI could potentially solve it much quicker, but the use of "human time to solve" is an attempt to create a metric that reveals long horizon complexity (as I understand it anyway). It's interesting because like the article notes, AI is really smashing benchmarks, but actual usefulness in automation of thought work is proving much more elusive. I think that collective experience of AI just not being that useful, or as useful as benchmarks suggest it should be, is captured in this metric.
- rishabhaiover 10mo agoI've practiced a healthy skepticism of the recent boom but I can't reason why the long horizon time wouldn't stretch to 8 hours or a week worth's of effort from next year. After Opus-4.5, governments and organizations should really figure out a path out of this storm because we're in it now.
- theptip 10mo agoDoubling time has been 7 months for a while, so you should expect 8h not 1 week next year.
- dwohnitmok 10mo agoIt's significantly accelerated to 4 months since the beginning of 2025, which puts 1 week within reach if things stay on trend. But yes 7 months is the more reliable long-term trend.
- pugio 10mo agoOpus looks like a big jump from the previous leader (GPT 5.1), but when you switch from "50%" to "80%", GPT 5.1 still leads by a good margin. I'm not sure if you can take much from this - perhaps "5.1 is more reliable at slightly shorter stuff, choose Opus if you're trying to push the frontier in task length".
- gizmodo59 10mo agoYeah. 50% of the time to throw away expensive tokens and limits is not ideal. But I bet by this time next year OSS models will be at that capability!
- Aperocky 10mo agoI think the problem here is LLM eventually pollute its context window with so much of the current task that the larger picture or architectural sanity is forgotten in favor of the current task at hand. And rarely is a software one and done, with a few round like this, the software architecture would have become schizophrenic. Combating this tendency usually require a lot of the work of these "long task" to be thrown away and more closely limiting what the AI is trying to do as they happen. The success of one "long task" is not necessarily a good thing!
- Leynos 10mo agoThis was why server-side compaction in GPT-5.2 was such a big deal. The model is by default provided with a tool that will prioritise the initial task and salient updates in context window compaction, and the new model has been trained to use it.
- karimQuant 10mo agoThe big issue is the 50%, if you switch to 80% it's much less. Now if you are in the wrong side of 50% given the task was 4hours. How much additional time to 4hours you need. repeat trying to get the task done 50%*50%->25% , 50%^4 -> 6.25%. the cost of bad luck is very high.
- bulbar 10mo agoIt's it bad luck though? I would've thought that if AI can't solve it first try the probability of fixing it in second try would be higher/lower (depending on the task).
- bentobean 10mo ago> We show that this metric has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months. If true, how much of this is a result of: 1. Genuine technical advancement or: 2. Shoveling trillions of dollars into compute resources in order to service incoming LLM requests in a way that is completely unrealistic over the long term? In other words… are we talking about genuine, sustainable innovation that we get to take with us moving forward and benefit from? Or are we talking about an “improvement” that is more akin to a mirage that will eventually disappear when the Ponzi scheme eventually collapses?
- dghost-dev 10mo agoGood point.
- emp17344 10mo agoI wonder how much of this stuff is attributable to true model advancement, or if it’s an improvement in the genetic harness? It’s impossible to separate strict model improvement from improvement in the associated tools.
- mediaman 10mo agoMuch of this is due to vastly better posttraining RL, not models that are much bigger. The idea that most of these gains comes from training really big models, or throwing immensely larger amounts of compute at it, is not really true.
- twotwotwo 10mo agoI'm conflicted about opining on models: no individual has actually done a large sample of real-world tasks with a lot of models to be able to speak with authority, but I kinda think we should each share our dubiously-informed opinions anyway because benchmarks aren't necessarily representative of real-world use and many can clearly be gamed. Anyhow, I noticed more of a difference trying Opus 4.5 compared to Sonnet 4.5 than I'd noticed from, for example, the last couple Sonnet bumps. Objectively, at 1.66x Sonnet's price instead of the old 5x, it's much more often practical to consider reaching for than past Opus models. Anthropic's basic monthly thing also covers a fair amount of futzing with it in CC. At the other extreme, another surprise of this family is that Haiku 4.5 with reasoning on is usable: better than Sonnet with thinking off according to some bencharks, and in any case subjectively decent for point edits, single-page thingies, and small tools.
- atleastoptimal 10mo agoThey should do a 95% and 99% version of the graphs, otherwise it's hard to ascertain whether the failure cases will remain in the elusive "stuff humans can do easily but LLM's trip up despite scaling"
- Davidzheng 10mo agoBig error bars and METR people are saying the longer end of the benchmark are less accurate right now. I think they mean this is a lower bound!
- scellus 10mo agoIt's complicated. Opus 4.5 is actually not that good at the 80% threshold but is above others at 50% threshold of completion. I read there's a single task around 16h that the model completed, and the broad CI comes from that. METR currently simply runs out of tasks at 10-20h, and as a result you have a small N and lots of uncertainty there. (They fit a logistic to the discrete 0/1 results to get the thresholds you see in the graph.) They need new tasks, then we'll know better.
- JohnnyMarcone 10mo agoThanks for this comment. I've been trying to find anything about the huge error bars. Do you have any sources you can share for further reading?
- iLoveOncall 10mo ago> current models have almost 100% success rate on tasks taking humans less than 4 minutes The contrary is easily verifiable by everyone individually. It's nowhere near 100%, or even 50% for few minutes tasks even with the best models in real world situations.
- ben_w 10mo agoI've only noticed that combination (failure of short everyday tasks from SOTA models) on image comprehension, not text. So some model will misclassify my American black nightshade* weeds as a tomato, but I get consistently OK results for text out from good models unless it's a trick question. * I recon, at least; looked like this to me: https://en.wikipedia.org/wiki/Solanum_americanum#/media/File:Solanum_americanum_(4898754585).jpg https://en.wikipedia.org/wiki/Solanum_americanum#/media/File...
- iLoveOncall 10mo agoThe research from Metr, and my comment, is exclusively related to software development tasks.
- ben_w 10mo agoRe-reading my comment, I realise I missed the most important part, the question. What examples can you give of "real world situations" where they fail? Obviously I don't want to use them for whatever that is.
- NiloCK 10mo agoI appreciate horizon expansion as a fundamental metric, but duration seems like too crude a measure. We used to like it when computers were fast. An infinitely unscrupulous model provider could double this five hour result by cutting your output tokens/second in half! This isn't only a question of gaming the metric: the very strong current small-fast models (4.5 Haiku, Gemini 3 Flash) have no hope of being measured fairly against this - they will succeed or fail much faster just because they are much faster. How about something like total output token count as the "long term horizon" metric instead?
- docstryder 10mo agoTask duration is the time it would take for humans to complete the task. The speed of the models and how how long they might take to complete the task is not part of this metric.
- scellus 10mo agoThe time (horizon) here is not that of the model completing the task, but a human completing the task.
- NiloCK 10mo agoWow that was a garbage comment! My introduction to this type of model measuring came from an interview where the repeatedly hammered-home point was that Sonnet 4.0 nailed a gigantic refactor (conversion of a large legacy asp.net or similar into react server-side components or similar) in a loop whose runtime was some large number of hours. I mistakenly attributed the same framing here.
- scotty79 10mo ago> As shown above, when we fit a similar trend to just the 2024 and 2025 data, this shortens the estimate of when AI can complete month-long tasks with 50% reliability by about 2.5 years. I don't think I have 50% success rate at month long tasks. Anything that exceeds one day is pretty hard.
- rich_sasha 10mo agoHow does "cost" per frontier task change with time? Extrapolating any exponential growth is always dangerous, but over say 3 years at this pace, we'd go from 2 hours to 70,or about 8 days' work. Quite scary. But what does cost do over the same timeline? Does it increase with computational complexity? Is it worse - because, IIRC, transformers computational cost is quadratic in context length. Is it better - some kind of economies of scale? I glanced thought the article but couldn't find any info on this.
- 0x000xca0xfe 10mo agoAfter spending many hours optimizing some routines I now think performance optimization is a great benchmark for identifiying how generally smart an AI is at helping with some specific piece of code. Solutions are quite easy to verify with differential testing and produce a number for direct comparison. Less code is usually better and you generally can't "cheat" by adding more cruft so it nullifies the additive bias. Good optimization requires significant understanding of the underlying structures. Everything has performance tradeoffs so it requires systemic thinking and not just stringing independent pieces together. So far I've found that Gemini Pro 3 was the best at reasoning about tricky SIMD code but the results with most models were pretty underwhelming.
- zkmon 10mo ago> We believe this work has important implications ... > First, our work demonstrates an approach ... The Conclusions section is not for making a sales pitch for your article. It is for summarizing any new knowledge the article brings out.
- yoan9224 10mo agoThe key insight from this benchmark is using "human-equivalent hours" rather than actual AI execution time. It's measuring capability complexity, not speed. What's interesting is the 50% vs 80% reliability gap. At 50% success rate on a 4-hour task, you're essentially gambling. If it fails, you've potentially wasted the 4 hours plus the time debugging why it failed. This is why I think the current "agent" paradigm needs human checkpoints at regular intervals. Let the AI work for 30 minutes, then review progress. Repeat. This way you catch drift early before it compounds. The other thing missing from these benchmarks: recovery ability. When the AI gets stuck on hour 3 of a 4-hour task, can it recognize the problem and backtrack? Or does it confidently continue down the wrong path?
- hnthrowaway121 10mo agoYou’ve only wasted the 4 hours if you didn’t spend them doing something else. At 50/50 it’s an ok bet if the debugging time is much less than the total human time, even if the loops are long, you might rather 4 hours of deep work on an important human thing or on just relaxing vs babysitting the LLM. Assuming that about half the time that will pay off with a correctly done thing with very little effort, it’s kind of amazing.
- afro88 10mo ago> The key insight from this benchmark is using "human-equivalent hours" rather than actual AI execution time. It's measuring capability complexity, not speed. > What's interesting is the 50% vs 80% reliability gap. At 50% success rate on a 4-hour task, you're essentially gambling. If it fails, you've potentially wasted the 4 hours plus the time debugging why it failed. Your first two paragraphs are at odds with each other. If it fails, you've potentially wasted the time it took the agent to *perform* the "it takes humans 4h" long task. Which in most cases is single digit minutes. That's why one of the solid use cases for agents is doing multiple throw away proof of concepts to explore a problem / new feature before deciding on a solution to actually implement. Usually you'd have time for one, or maybe none. If it fails you've lost a maybe 10 minutes, but likely learned something new about the potential solution.
- bicepjai 10mo agoIMHO, in the software field, learning can be simpler to 2 phases. The first one is exploration, where we read blogs, docs, and books; listen to lectures and talks. Then comes the second phase of exploitation, where we actually use the thing we learned. You can think of all those “learning from scratch” videos as someone who is doing the phase 2. I love the phase one and most of the time don’t have time and energy to sit down and go through the phase 2. Nowadays, I feel like the 2 phases are combined, thanks to LLMs. For instance, I wanted to do some animation for visualizations. This week, I learned AnimeJS by watching CCAgent create the animation I wanted, which was interspersed with questions that were answered with diagrams and text, which accomplishes the phase 1. I do not like letting them run the show. Then comes phase 2, where I organize the code, abstract things, rewrite code, still use their help for long rewrites, but totally my ideas and mine only. This saves time tremendously.
- yoan9224 10mo agoThe key insight from this benchmark is using "human-equivalent hours" rather than actual AI execution time. It's measuring capability complexity, not speed. What's interesting is the 50% vs 80% reliability gap. At 50% success rate on a 4-hour task, you're essentially gambling. If it fails, you've potentially wasted the 4 hours plus the time debugging why it failed. This is why I think the current "agent" paradigm needs human checkpoints at regular intervals. Let the AI work for 30 minutes, then review progress. Repeat. This way you catch drift early before it compounds. The other thing missing from these benchmarks: recovery ability. When the AI gets stuck on hour 3 of a 4-hour task, can it recognize the problem and backtrack? Or does it confidently continue down the wrong path?
- dvfjsdhgfv 10mo ago> This is why I think the current "agent" paradigm needs human checkpoints at regular intervals. Let the AI work for 30 minutes, then review progress. Repeat. This way you catch drift early before it compounds. The problem with this approach is that in 30 minutes, an agent is able to produce a massive amount of stuff. Reviewing all this is a nightmare, in the sense that on the surface it seems fine and it often works, until it doesn't. The bugs introduced are often subtle and their effects manifest later, if ever. So, for stuff that matters (to me), I prefer not to use agents at all. Maybe things will change in a year, or 5, or 10. I will be giving it a try. but for the moment it's just not worth it, and the upside-down workflow it pushes on me is just making me tired and lose satisfaction from doing my job.
- mkoubaa 10mo agoAsk not what the agent can do you for you, ask what you can do for the agent. If you fail to break up the task into agent sized chunks, you're the problem.
- sshh12 10mo agoFor folks interested in some of the nuances of this benchmark, I just posted this deep dive: https://blog.sshh.io/p/understanding-ai-benchmarks https://blog.sshh.io/p/understanding-ai-benchmarks
- big-chungus4 10mo ago"Train adversarially robust image model" is not a long task imo
- leecommamichael 10mo agoI read their citations (which are actually the same authors of this paper) and they also define using Python's built-in web server to "build a web server" as a long task.