9 ms·
What Happens When the Cost of Intelligence Drops 100x
- bkd9 1mo agoAuthor here. I made these plots because I had been searching for them for months and never found quite what I wanted: how the cheapest way to reach a fixed capability level has moved over time. Artificial Analysis publishes enough data to reconstruct it. If someone knows of a source that already tracks this, with historical prices, please share.
- andai 1mo agoThanks for the effort you put into this. I wonder if it might drive the point even further if the graph scales were linear? Or maybe the progress has been so great that this would make the graphs unreadable? https://xkcd.com/1162 https://xkcd.com/1162
- brabel 1mo agoIt blows my mind how fast models are getting better and this is the first article I’ve seen that shows just that and leaves almost no room for disagreement. Well done.
- jnwatson 1mo agoThe animated Pareto front graph by time is truly inspired. Well done!
- AnotherGoodName 1mo agoI think a big one is robotics. A robot can today fold your laundry. It takes ~10mins per item. Seriously. It takes a long time to process the image find the corner move the claw to the corner of the shirt and attempt to straighten before folding. Robots right now generally move at glacial speeds. You might have seen robots doing flips in semi controlled environments but watch how slowly they open doors etc. processing time is a major bottleneck.
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- andai 1mo agoWait til the robots get on Cerebras, it'll set your pants on fire.
- LoganDark 1mo agoOnly if the fire manages to escape your wallet!
- segmondy 1mo agoYou must not have been paying attention to development with robots, there are many videos of robots moving really fast in "non controlled environments"
- airstrike 1mo agoYes, I think I saw one in a video titled "Robocop"
- TheAceOfHearts 1mo agoHonestly, most of the videos I've seen of robots moving around quickly aren't actually doing anything useful. We've had really impressive tech demos for the past 15 years of robots dancing and jumping around. But I don't want a dancing robot, I want a robot to make me a BLT, wash my dishes, take out the trash, and fold my laundry. The most recent video which actually impressed me was a demonstration from Gemini Robotics 2, where a robot was shown autonomously removing the bag from a trash can and folding the loops closed in real time. I don't follow robotics advances closely so it's possible I'm just ignorant, do you know any autonomous robotics demonstrations of useful activities that you would suggest checking out?
- infecto 1mo ago
- brotchie 1mo agoThere’s still 50-500x cost reduction in “this is only an engineering problem” low hanging fruit from specialized chips to run the models + improved distillation. Entirely feasible that by 2031, Fable 5 (or greater) intelligence level models will run cool on smart phones, if not sooner.
- andai 1mo agoI saw a 1B model yesterday that was fine tuned on Fable output. I found that hilarious, but it did actually make all the scores go up. (Actually talking to it, it was about as coherent as you'd expect, i.e. 3/10) The floor for "actually usable model" keeps dropping though. (Seems to be about 27B right now?)
- tgv 1mo agoYou're betting on getting getting ridiculously powerful chips to run on batteries in a tiny housing without cooling, while we can't even get enough RAM? It would be a terrible waste of resources. Now we already have TFLOPs wasting in our pockets and backpacks, then we'll have PFLOPs idling, because there's so much time between prompts. Much more efficient to batch it on a server.
- kaashif 1mo agoAt some point we'll have enough RAM, surely. The incentives to produce more are huge and all the fabs are booked out. Maybe it'll take 10 years or 20 years. <5 years is not long enough for manufacturing to catch up. Not much of a comment on the phone stuff but I'd caution against suggesting technology will never be good enough to do X. Maybe it'll be horrendously wasteful but it might happen.
- missedthecue 1mo agoIf fable 5 could eventually run on a smartphone, I wonder what we'll get out of datacenters. Musk and others are trying to build 100gw of compute by 2030. Will we get something 1-10 million times better than Fable 5, or will the parity gap between local and datacenter capability pair down enormously?
- qsera 1mo agoIdiocy becomes rampant!
- andai 1mo agoAGI ≈ Ralph × Infinite persistence Just like real life!
- perching_aix 1mo agobecomes rampant? Where have you been the past decades? Hell, centuries and millenia? If we're going cynical, may as well go full throttle.
- Flamkuchlo 1mo agoWhats your problem? Even if its not intelligence, a LLM found a bug due to one error message, fixed it, created a PR and it solved it. If an LLM is only able to do all of this after training on it and never achieving AGI, we already at the point were it is cheaper to teach one LLM one problem than teaching humans to do so.
- promptsphere 1mo ago[flagged]
- MichaelNolan 1mo ago100x seems like an underestimate. Even with no model improvements, we should see that sort of reduction. Looking at TSMC’s margins, Nvidia’s margins, and OAI/Anth (alleged) margins on inference, there is a room for a 100x reduction. Right now all three of those are at abnormally high levels. Competition will come for all three.
- andai 1mo agoA year ago I had an aha moment, when I realized that for my purposes, Gemini Flash was not only 9x cheaper, but 3x faster than Gemini Pro, while producing identical output. Who's the best model now! For a lot of tasks, even small models have saturated them a while ago, and then going cheaper and faster is just pure gains. For coding I also prefer to do it interactive/realtime, micro-prompting, surgical edits, which the small models can handle just fine. And then at the top, the real question is consistency. Not "can they do it" but "reliably enough that you don't need to constantly double check everything." (In my experience, not quite there yet, although it's getting way better.)
- bdhdhduuyd 1mo agoPersonally I still see LLMs as very advanced search engines which lack intelligence. To me it seems that the cost of getting data is reduced by LLMs, not the cost of intelligence. I mean: we tell the model what we want to achieve, and the model responds with the right data in de form of code in seconds. That's why 'stackoverflow programmers' will have a hard time competing with LLMs but engineers are still needed for their intelligence. Well that's just my 2 cents.
- dboreham 1mo agoI think this viewpoint fails to understand what "intelligence" is. The idea must be that intelligence is some special thing that only humans have. So when machines couldn't do jack s... we said "it's the Turing test". When machines blew through the Turing test we said "that was just prediction..not really intelligence, that's different". It's not different. The delusion humans have is that intelligence is special and magical. It's not. It's just nature's prediction machine. A very fancy version to be sure. But not qualitatively different . All statements that "oh but it'll never be able to do that" will prove false.
- perching_aix 1mo agoI notice that a lot of these terms are squarely humanist for certain people, so any kind of allegation that machines are exhibiting them as traits will be a complete showstopper for them. You'll either be strung along on an infinite goalpost moving exercise, or be accused of either anthropomorphization [0] in the kinder cases, or straight up mental illness in the less so kind cases. Never will they stop to consider that maybe you're simply working with a post-humanist understanding of these words, as that basically doesn't make sense to them, and as they are usually quite vested for it to stay that way. I remember in one of the Hugging Face incident threads here, simply acknowledging that the agents were operating autonomously was super controversial. Thousands of years old concept [1], still inherently human for a lot of people. To be clear, I'm not trying to be judgemental with this, I more consider it to be a communications breakdown, and find that to be frustrating instead. I'm not really sure how to meaningfully help it either, cause no matter how one slices it, you will in the end ask these people do desecrate these terminologies in favor of some more twisted-seeming ones. Same the other way around, the humanist understanding of these terms is basically non-workable. [0] as opposed to personification, which is what people are actually doing almost always: https://en.wikipedia.org/wiki/Personification https://en.wikipedia.org/wiki/Personification [1] https://en.wikipedia.org/wiki/Automaton https://en.wikipedia.org/wiki/Automaton
- Multiplayer 1mo agoThis means a great de-risking is happening for the costs of deploying somewhat autonomous agents. This has profound implications on the timeline of deployable personal agents. Cost was a significant factor for many people during the OpenClaw frenzy, specifically when they let their agents run somewhat wild. It will become much more palatable, or already has, to install whatever the next generation of token consuming autonomous systems will be.
- smallnix 1mo ago[dead]
- Balooga 1mo agoJevons Paradox [1] > when technological improvements that increase the efficiency of a resource's use lead to a rise, rather than a fall, in total consumption of that resource. [1] - https://en.wikipedia.org/wiki/Jevons_paradox https://en.wikipedia.org/wiki/Jevons_paradox Las Vegas replaced the expensive incandescent lighting on the strip with cheaper to run LED equivalents. But the costs didn't come down because they were able to add more lights and larger displays. I think the same will happen with tokens. As the cost of tokens comes down, these models will just consume more tokens.
- perching_aix 1mo agoSounds like a variation on the induced demand principle: https://en.wikipedia.org/wiki/Induced_demand https://en.wikipedia.org/wiki/Induced_demand Related: - Parkinson's law: "Work expands to fill the available time." https://en.wikipedia.org/w/index.php?title=Parkinson%27s_Law https://en.wikipedia.org/w/index.php?title=Parkinson%27s_Law - Lewis–Mogridge position: "Traffic expands to meet the available road space." https://en.wikipedia.org/wiki/Lewis%E2%80%93Mogridge_position https://en.wikipedia.org/wiki/Lewis%E2%80%93Mogridge_positio... And I pretty much just plain agree, this is exactly what will happen. I don't think there's anything wrong with it (in isolation) either, though I do already find myself pointing out that we're misusing LLMs at work sometimes (most notably, a recent mini project could have been a jinja template - and it did become one thanks to me pushing back on this). Abundance is one thing, waste and misuse is another.
- bena 1mo agoAlso Jevon's Paradox: the more efficient a system becomes, the more it gets used, making it use more resources rather than less.
- AndrewKemendo 1mo agoTurns out Grey Goo was just thermal paste
- euroderf 1mo ago
- andai 1mo ago> Reading everything becomes the default. At a cent per document, a model can read every paper I love how in our day "reading everything" means "the computer reads it for me". I expect soon the computer will be able to go on bicycle rides, and spend time with my wife.
- sssilver 1mo agoHasn't the computer already been spending time with your wife?
- lubujackson 1mo agoYou're absolutely right!
- smugtrain 1mo agoAsymmetric multiprocessing with your mom
- iririririr 1mo agoit's called Instagram
- euroderf 1mo ago"We're living in an attention economy." "So's your mom."
- bellowsgulch 1mo agoFuturama did it first!
- gs17 1mo ago"Robot, experience this tragic irony for me!"
- NooneAtAll3 1mo agonooooooooooo
- nchmy 1mo agoI've been working with the chinese open models for 4 months. They are more than capable for a tiny fraction of the cost of the frontier ones. And yet they also continue to get significantly better and (Deepseek's recent price increase aside) cheaper. Its hard to fathom how the truly frontier stuff will be able to compete long-term.
- jostmey 1mo agoAnd why won’t the frontier models continue to become better? The open models are getting better but so are the frontier models. The frontier models might remain in a constant race to remain ahead
- ForHackernews 1mo agoThey are running out of novel, clean training data and compute. There is probably a limit to how much improvement can be squeezed out of LLMs. Recent improvements have been more about orchestration and "reasoning" loops (i.e. iteratively feeding context back through the model).
- svachalek 1mo agoFor base models they really must be running out of new training materials. It's more about size and architecture. But they still seem to be getting big strides out of improving the post training. They keep dropping point upgrades in under 8 weeks lately, which is an insane pace for product release. At some point it's going to slow down but we're not close yet imo.
- HappyPanacea 1mo agoDiminishing returns on both intelligence and training, mostly
- sweetjuly 1mo agoI don't think it's a matter of "staying ahead"; the proprietary frontier models are better, but the trouble for them is that open weight models are good enough in increasingly many cases. This raises the floor on the frontier companies and cuts their total addressable market by commodifying the easier LLM tasks. This is really the central argument of tfa :)
- LordHumungous 1mo ago> Reading everything becomes the default. At a cent per document, a model can read every paper in a field, every record in an archive, every email, or every message in a support queue as a matter of routine Pretty much already happening
- newAccount2025 1mo agoI’m loving small models. The gemma4 26/31b models have been deeply impressive on weird prose analysis tasks that I am working on. Nova-micro is really stupid but is extremely fast when it’s smart enough to do something. I’m trying to be disciplined about able to evaluate quality vs cost everywhere for real systems built on this stuff. I probably need to get off Bedrock because it’s missing a lot of other little models that might be good competitors.
- deleted 1mo ago[deleted]
- jbotdev 1mo agoI think speed is actually going to be a bigger factor than cost. Even projects where “money is no object” often hit a wall with LLM response times. Sure you can speed things up with parallel work under subagents, but as with parallelizing traditional computational tasks, there are diminishing gains. I keep hearing people saying just change the way you work to trust long-running agents and multi-task more, because they’re too slow to work with interactively for many use cases. I think that’s painful in a world where we expect humans to still heavily guide and interact with agents for their day-to-day work.
- perching_aix 1mo agoGiven the 750 tok/sec GPT 5.6 Sol Ultrafast (via Cerebras), the many-1000 tok/sec Chinese models, and the 15000 tok/sec Taalas HC1, I think we're well on the way towards seeing that solved too. Combine the two, and yeah, wild ride incoming. What's especially bewildering to me is that translated back to raw bandwidth, even 15000 tok/sec is just like what, 75 KB/s? Extremely meager amounts of data, moving mountains. It's already kinda funny seeing LLMs throw out effort estimates in wall time terms. It's always some "hours, days, weeks" tier thing, when in reality, it's gone and done in minutes.
- cactusplant7374 1mo agoOne of my projects has an estimate of 3,000+ hours. It seems accurate. It has spent months working on it.
- perching_aix 1mo agoAround the clock?
- cactusplant7374 1mo agoYes, with a few exceptions. It's much easier now that the five hour limit has been removed from codex.
- 1mo ago
- vanuatu 1mo agoI think what a lot of people miss about jevon's paradox is the elasticity of demand of the underlying resource textiles had jevons paradox, and many more textile workers were employed even when textile machines were being created, until we saturated the demand for cheap clothing in the world and then textile workers were kaput (same for farming, and horses) software is currently undergoing jevons paradox, but it's very unknown how high the ceiling of demand for software is. web dev might be doomed, but software in general i think is probably limitless Intelligence is also probably unbounded (atm software and intelligence are very closely tied together). its very possible token spend rides up the curve forever.
- quantified 1mo agoBrings up the question of what the intelligence is used for. Humans exploited intelligence for competition. With each other to wipe out other Homo species, mate more and collect resources, with other animals to limit predator impact and gain food. Intelligence will be used offensively by corporations and their people to extract more from consumers (make pricing opaque, terms of service more complicated, etc.) and scams far more sophisticated. The "consumer" will need extra intelligence to fight all that off. There are only so many meals you can expertly produce, shirts to fold, itineraries to fun places you can execute, but there's a practical infinity of traps to set and avoid.
- danieltk76 1mo agothis guy gets it
- imnotr0b0t 1mo agoThe article is solid. But there is a nuance what he skipped — quality vs price. Sure, GPT-5.6 Luna for pennies can do the same thing what Claude 4.5 Sonnet did for a dollar a year ago. Except Sonnet back then actually carried the codebase, while Luna... eh, not so much. And another thing, speed. You can make it cheaper as much as you want, but if a model thinks for half a minute you save cents but lose time.
- Tade0 1mo agoThe other day I saw a benchmark of coding models that fit in 8GB VRAM - for reference some version of Mistral was added, normally requiring 32GB, but moving along at 4-5tok/sec when partially offloaded to CPU. Surprisingly, some of the small models would not only give worse results, but also took longer than Mistral, because they were thinking so much. That is an important detail which I was previously overlooking.
- deleted 1mo ago[deleted]
- cush 1mo agoIt’s way too early to tell the true cost of intelligence. Inference is still heavily subsidized, and training is apparently being funded by a mountain of free money
- infecto 1mo agoHow do we know inference is heavily subsidized? Are all these third party Chinese model providers subsidizing the true cost?
- cush 1mo ago> How do we know inference is heavily subsidized? Balance sheets > Are all these third party Chinese model providers subsidizing the true cost? It could be argued the subidies are even heavier than American model providers as the price war among Chinese providers is so fierce. https://www.scmp.com/tech/big-tech/article/3358868/after-triggering-price-war-deepseek-reverses-course-surcharge-peak-hour-api-use https://www.scmp.com/tech/big-tech/article/3358868/after-tri...
- infecto 1mo agoYou did not answer the question at all. Third party inference providers for Chinese models are in the US and there are so many players it would be shocking they are all heavily subsidizing it.
- cush 1mo agoIt would be shocking to you if in a market with many competing sellers would there be price competition such that those sellers are operating at a loss? If you’re trying to get me to argue that every seller is operating at a loss, I don’t know. But overall the market absolutely is, again because balance sheets
- infecto 1mo agoBalance sheets is not an answer. Sorry I am just asking for material proof. It would indeed be pretty shocking for every market participant, large or small, funded or unfunded to be operating inference at a loss. Plans are definitely subsidized. Inference (token consumption) has never been proven to be operating at a loss and now that you can run many of the large SOTA models from China on sparks or other clustered servers you can figure out some of the math. Not trying to argue but just saying “balance sheets” makes zero sense. Yes lots of capex spend, hard to say if anyone has spent too much. At the same time demand is increasing for compute.
- fabsalvadori 1mo agoThere is a second-order effect I see, here: when generating work becomes dramatically cheaper, failed work also becomes dramatically more common. With coding agents, cheap intelligence doesn't just mean producing the same software for less, but rather trying five implementations, letting agents run longer, touching larger scopes, and supervising fewer intermediate steps. That changes which infrastructure matters. When attempts are expensive, you optimize for success, while when attempts are cheap, you start optimizing for how cheaply you can inspect, reject and recover from failure. Git was an enormous enabler of cheap human experimentation, and I suspect we'll end up building analogous primitives around autonomous work.
- dominotw 1mo agonothing because this shit is not "intelligence" ppl are doing all sorts of gymnastics to tell claude to slow its roll with verbosity.
- iririririr 1mo agoall the optimistic pundits ignore the blatant second order consequence of this: the entire economy is pegged on this NOT being cheap! all the US economy is tied to video cards being used in lieu of gold. cost dropping 100x means the economy bottom falls out.
- maxglute 1mo agoI imagine some form of token consensus maxing. Repeat tasks X times concurrently to find agreement or outliers. Millions of token burnt deciding what sock to wear.
- hulitu 1mo ago> What Happens When the Cost of Intelligence Drops 100x Yes, what ? The "Cost of Intelligence" has risen due to AI. Earlier, we only had our mistakes to correct. Now we have also AI's mistakes to correct.