4 ms·
The METR study cited here is very interesting. "In the METR study, developers predicted AI would make them 24% faster before starting. After finishing 19% slow
by fancyfredbot 9mo ago
The METR study cited here is very interesting.
"In the METR study, developers predicted AI would make them 24% faster before starting. After finishing 19% slower, they still believed they'd been 20% faster."
I hadn't heard of this study before. Seems like it's been mentioned on HN before but not got much traction.
- Sharlin 9mo agoPlenty of people have been (too) quick to dismiss that study as not generally applicable because it was about highly experienced OSS devs rather than your average corporation programmer drone.
- fancyfredbot 9mo agoThat's interesting context for sure, but the fact these were experienced developers makes it all the more surprising that they didn't realise the LLM slowed them down.
- Sharlin 9mo agoMeasuring programming productivity in general is notoriously difficult, subjectively measuring your own programming productivity is even worse. A magic LoC machine saying brrrrrt gives an overoptimistic sense of getting things done.
- _aavaa_ 9mo agoThe issue I have with the paper is that it seems (based on my skimming) that they did not pick developers who were already versed with AI tooling. So they're comparing (experienced dev working in the way they're comfortable) vs (experienced dev working with new tool for the first time and not having passed the productivity slump from onboarding).
- Sharlin 9mo agoLongitudinal studies are definitely needed, but of course at the time the research for this paper was done there weren't any programmers experienced with AI assist out there yet.
- pydry 9mo agoThe thing I find interesting is that there is trillions of dollars in valuations hinging upon this question and yet the appetite to spend a little bit of money to repeat this study and then release the results publicly is apparently very low. It reminds me of global warming where on one side of the debate there some scientists with very little money running experiments and on the other side there were some ridiculously wealthy corporations publicly poking holes in those experiments but who secretly knew they were valid since the 1960s.
- Terr_ 9mo agoYeah, it's kind of a Bayesian probability thing, where the impressiveness of either outcome depends on what we expected to happen by default. 1. There are bajillions of dollars in incentives for a study declaring "Insane Improvements", so we should expect a bunch to finish being funded, launched, and released... Yet we don't see many. 2. There is comparatively no money (and little fame) behind a study saying "This Is Hot Air", so even a few seem significant.
- simonw 9mo agoI see it brought up almost every week! It's a firm favorite of the "LLMs don't actually help write code" contingent, probably because there are very few other credible studies they can point to in support of their position. Most people who cite it clearly didn't read as far as the table where METR themselves say: > We do not provide evidence that: > 1) AI systems do not currently speed up many or most software developers. Clarification: We do not claim that our developers or repositories represent a majority or plurality of software development work > 2) AI systems do not speed up individuals or groups in domains other than software development. Clarification: We only study software development > 3) AI systems in the near future will not speed up developers in our exact setting. Clarification: Progress is difficult to predict, and there has been substantial AI progress over the past five years [3] > 4) There are not ways of using existing AI systems more effectively to achieve positive speedup in our exact setting. Clarification: Cursor does not sample many tokens from LLMs, it may not use optimal prompting/scaffolding, and domain/repository-specific training/finetuning/few-shot learning could yield positive speedup https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...
- fancyfredbot 9mo agoWeird, you shouldn't really need to list the things your study doesn't prove! I guess they anticipated that the study might be misrepresented and wanted to get ahead of that. Their study still shows something interesting, and quite surprising. But if you choose to extrapolate from this specific setting and say coding assistants don't work in general then that's not scientific and you need to be careful. I think the studyshould probably decrease your prior that AI assistants actually speed up development, even if developers using AI tell you otherwise. The fact it feels faster when it is slower is super interesting.
- simonw 9mo agoThe lesson I took from the study is that developers are terrible at estimating their own productivity based on a new tool. Being armed with that knowledge is useful when thinking about my own productivity, as I know that there's a risk of me over-estimating the impact of this stuff. But then I look at https://github.com/simonw https://github.com/simonw which currently lists 530 commits over 46 repositories for the month of December, which is the month I started using Opus 4.5 in Claude Code. That looks pretty credible to me!
- MagicMoonlight 9mo agoI can believe it. It will zero-shot a full system for you in 5 minutes, but then if you ask for a minor change to that system it will completely shit the bed. And you have no understanding of what it has written, so you’d have to check everything.