8 ms·
Are we plotting against cost? How is the capability advancement vs dollars paid for development? By my read of the (very sparse) data, we're getting linear im
by oudlys 4mo ago
Are we plotting against cost? How is the capability advancement vs dollars paid for development?
By my read of the (very sparse) data, we're getting linear improvements in capability for super-linear increases in costs. [1] Indicates that by 2027 models will cost $1 billon to train. Dario estimates that model runs will cost $10 billion in 2026 [2]. That to me indicates costs are potentially growing faster than capability. Maybe by quite a bit.
If the value prop of LLMs doesn't prove out, that won't last. I'm of the opinion there is no data that shows actual economic value being delivered by models. The best data shows that LLM use might be destroying value [3].
[1] https://epoch.ai/publications/how-much-does-it-cost-to-train-frontier-ai-models https://epoch.ai/publications/how-much-does-it-cost-to-train...
[2] https://lexfridman.com/dario-amodei-transcript/ https://lexfridman.com/dario-amodei-transcript/
[3] https://unessays.substack.com/p/talk-is-cheap https://unessays.substack.com/p/talk-is-cheap
- simianwords 4mo ago>By my read of the (very sparse) data, we're getting linear improvements in capability for super-linear increases in costs. [1] Indicates that by 2027 models will cost $1 billon to train. Dario estimates that model runs will cost $10 billion in 2026 [2]. That to me indicates costs are potentially growing faster than capability. Maybe by quite a bit. This is true and well established. As long as you get any improvement whatsoever, it is worth spending to train since it pays off during. Imagine training was not $1 billion but $100 billion but the performance improved by just 10%. This is still worth it because you can squeeze out the profits across years and years right? The improvement is ever lasting. > The best data shows that LLM use might be destroying value [3]. This is basically a conspiracy theory and if you really believed this, you should not have led with "How is the capability advancement vs dollars paid for development?" because if there were no value, it doesn't really matter how much you invest.
- oudlys 4mo ago>This is basically a conspiracy theory I think this is pretty uncharitable, especially when I've provided you with a dataset you can evaluate yourself and an argument you can review for logical inconsistency. I have worked quite hard to locate data that supports your thesis, I can't find it. I've at least gone to the effort of documenting that search. Before you throw around such strong convictions, I suggest you actually look for yourself.
- simianwords 4mo agoRespectfully, your link is not very convincing. But what’s interesting is that you are commenting on a post where Dario is suggesting that LLMs are so extremely powerful that they can take over, help synthesise bioweapons, help in warfare, help in drug discovery — the whole post here is to try and regulate this. If you believe AI can’t even create positive value let alone discover new things then your problem is somewhere else and not in something like “but training costs a lot”. So it is absolutely strange and contrasting to see you believe that LLMs are so weak as to create negative value while the CEO is asking about regulations because AI is too powerful. I don’t think I can convince you that AI is actually that powerful. But let me ask you something directly: if you believe what you believe, you should also acknowledge that AI doesn’t need regulations in the context Dario is proposing since obviously AI can’t do anything he predicts. Do you agree?
- deleted 4mo ago[deleted]
- dag100 4mo ago> So it is absolutely strange and contrasting to see you believe that LLMs are so weak as to create negative value while the CEO is asking about regulations because AI is too powerful. You wouldn't ask a chemistry professor to write code. So just because LLMs create negative value for software development doesn't mean that they can't be helpful for bioweapons synthesis, especially considering the range of chemistry and biology sources Anthropic would have fed to its LLM that wouldn't be publicly accessible. The LLM doesn't even need to be particularly accurate so long as the amateur bioweapons researcher takes adequate precautions before following its instructions and does some background research beforehand.
- simianwords 4mo agoThis is a ridiculous stance to take. That LLMs are simultaneously negative value but can also help synthesise bioweapons. It’s the sort of stance you take when you already feel ideologically against AI. I don’t think it’s coherent.
- aspenmartin 4mo agoI appreciate the data here but I don't think the read is quite right; Saying we have linear capability for super-linear cost compares an unbounded variable (dollars) to bounded instruments (because benchmarks saturate). On unbounded measures, growth is exponential; you can see METR time horizons double every ~4-7 months (https://metr.org/blog/2026-1-29-time-horizon-1-1/ https://metr.org/blog/2026-1-29-time-horizon-1-1/). And capability being proportional to log(compute) is what the scaling law predicts. Epoch puts training cost growth at ~2.4x/year as your link shows. Meanwhile cost for fixed capability falls ~10-40x/year (https://epoch.ai/data-insights/llm-inference-price-trends https://epoch.ai/data-insights/llm-inference-price-trends), and lab revenue is growing ~10x/year! Anthropic went from $1B to $9B to $30B+ run rate in ~15 months, OpenAI ~$25B. On [3]: the "destroying value" conclusion flips sign on an assumed 15% baseline rework rate. The report's most direct metric is +16% merged PRs per dev. The RCT evidence is genuinely mixed (METR: -19%, with n = 20 and Claude 3.x; Cui et al: +26%) but its just super hard to do this well, I think Faros stuff was pretty cool, I haven't seen this before so thank you for the reference.
- balefulboy 4mo agoMETR's time horizon is not a reliable metric of LLM capability growth: https://www.transformernews.ai/p/against-the-metr-graph-coding-capabilities-software-jobs-task-ai https://www.transformernews.ai/p/against-the-metr-graph-codi...
- oudlys 4mo agoWow. This deserves to be much more widely read. Thank you for this.
- aspenmartin 4mo agoYes I've seen this before, and while the critiques are fair and high quality (and unfortunately not unique to METR) we're missing the forest for the trees here. First of all, if you take the articles critiques and work out the implications on the METR graph, all you're doing is shifting the curve up or down, it doesn't change the fact that progress is scaling exponentially. While it is technically possible the universe could be throwing a massive pathological curveball to change the conclusion from METR data (which is we've been seeing exponential growth over the last 6 years), I think that seems very far from likely. The fact that we see the same behavior from a variety of sources over a wide variety of tasks and domains is a pretty clear indication that METR while certainly far from perfect is actually painting a consistent picture at least in terms of the rate of progress. You can look at ECI for a summary benchmark statistic, which does NOT use METR's benchmark, and you see a similar trend. Same with SWE-bench where the task distribution is far more in domain for real world problems. It is a bummer that this METR data can't be better funded. It would probably take $1M or so to really beef it up properly which any of these labs probably have in their couch cushions.