21 ms·
The compute moat is getting absolutely insane. We're basically at the point where you need a small country's GDP just to stay in the game for one more generatio
by llamasushi 1y ago
The compute moat is getting absolutely insane. We're basically at the point where you need a small country's GDP just to stay in the game for one more generation of models.
What gets me is that this isn't even a software moat anymore - it's literally just whoever can get their hands on enough GPUs and power infrastructure. TSMC and the power companies are the real kingmakers here. You can have all the talent in the world but if you can't get 100k H100s and a dedicated power plant, you're out.
Wonder how much of this $13B is just prepaying for compute vs actual opex. If it's mostly compute, we're watching something weird happen - like the privatization of Manhattan Project-scale infrastructure. Except instead of enriching uranium we're computing gradient descents lol
The wildest part is we might look back at this as cheap. GPT-4 training was what, $100M? GPT-5/Opus-4 class probably $1B+? At this rate GPT-7 will need its own sovereign wealth fund
- duxup 1y agoIt's not clear to me that each new generation of models is going to be "that" much better vs cost. Anecdotally moving from model to model I'm not seeing huge changes in many use cases. I can just pick an older model and often I can't tell the difference... Video seems to be moving forward fast from what I can tell, but it sounds like the back end cost of compute there is skyrocketing with it raising other questions.
- ljlolel 1y agoThe scaling laws already predict diminishing in returns
- renegade-otter 1y agoWe do seem to be hitting the top of the curve of diminishing returns. Forget AGI - they need a performance breakthrough in order to stop shoveling money into this cash furnace.
- jayde2767 1y ago"cash furnace", so aptly put.
- general1465 1y agoYep we do. There is a 1 year old video on YouTube, which describes this limitation https://www.youtube.com/watch?v=5eqRuVp65eY https://www.youtube.com/watch?v=5eqRuVp65eY Called efficient compute frontier
- fredoliveira 1y agoI think that the performance unlock from ramping up RL (RLVR specifically) is not fully priced into the current generation yet. Could be wrong, and people closer to the metal will know better, but people I talk to still feel optimistic about the next couple of years.
- duxup 1y ago>cash furnace They don't even burn it on on AI all the time either: https://openai.com/sam-and-jony/ https://openai.com/sam-and-jony/
- dmbche 1y ago"May 21, 2025 This is an extraordinary moment. Computers are now seeing, thinking and understanding. Despite this unprecedented capability, our experience remains shaped by traditional products and interfaces." I don't even want to learn about them every line is so exhausting
- duxup 1y agoAgreed, that whole page is brutal to read.
- serf 1y agoI was expecting a wedding or birth announcement from that picture framing and title. "We would like to introduce you to the spawn of Johnny Ive and Sam Altman, we're naming him Damien Thorn."
- mikestorrent 1y agoInference performance per watt is continuing to improve, so even if we hit the peak of what LLM technology can scale to, we'll see tokens per second, per dollar, and per watt continue to improve for a long time yet. I don't think we're hitting peak of what LLMs can do, at all, yet. Raw performance for one-shot responses, maybe; but there's a ton of room to improve "frameworks of thought", which are what agents and other LLM based workflows are best conceptualized as. The real question in my mind is whether we will continue to see really good open-source model releases for people to run on their own hardware, or if the companies will become increasingly proprietary as their revenue becomes more clearly tied up in selling inference as a service vs. raising massive amounts of money to pursue AGI.
- ethbr1 1y agoMy guess would be that it parallels other backend software revolutions. Initially, first party proprietary solutions are in front. Then, as the second-party ecosystem matures, they build on highest-performance proprietary solutions. Then, as second parties monetize, they begin switching to OSS/commodity solutions to lower COGS. And with wider use, these begin to outcompete proprietary solutions on ergonomics and stability (even if not absolute performance). While Anthropic and OpenAi are incinerating money, why not build on their platforms? As soon as they stop, scales tilt towards an apache/nginx type commoditized backend.
- reissbaker 1y agoAccording to Dario, each model line has generally been profitable: i.e. $200MM to train a model that makes $1B in profit over its lifetime. But, since each model has been more and more expensive to train, they keep needing to raise more money to train the next generation of model, and the company balance sheet looks negative: i.e. they spent more this year than last (since the training cost for model N+1 is higher), and the model this year made less money this year than they spent (even if the model generation itself was profitable, model N isn't profitable enough to train model N+1 without raising — and spending — more money). That's still a pretty good deal for an investor: if I give you $15B, you will probably make a lot more than $15B with it. But it does raise questions about when it will simply become infeasible to train the subsequent model generation due to the costs going up so much (even if, in all likelihood, that model would eventually turn a profit).
- viscanti 1y agoWell how much of it is correlation vs causation. Does the next generation of model unlock another 10x usage? Or was Claude 3 "good enough" that it got traction from early adopters and Claude 4 is "good enough" that it's getting a lot of mid/late adopters using it for this generation? Presumably competitors get better and at cheaper prices (Anthropic charges a premium per token currently) as well.
- dom96 1y ago> if I give you $15B, you will probably make a lot more than $15B with it "probably" is the key word here, this feels like a ponzi scheme to me. What happens when the next model isn't a big enough jump over the last one to repay the investment? It seems like this already happened with GPT-5. They've hit a wall, so how can they be confident enough to invest ever more money into this?
- bcrosby95 1y agoI think you're really bending over backwards to make this company seem non viable. If model training has truly turned out to be profitable at the end of each cycle, then this company is going to make money hand over fist, and investing money to out compete the competition is the right thing to do. Most mega corps started out wildly unprofitable due to investing into the core business... until they aren't. It's almost as if people forget the days of Facebook being seen as continually unprofitable. This is how basically all huge tech companies you know today started.
- yieldcrv 1y agoLocally run video models that are just as good as today’s closed models are going to be the watershed moment The companies doing foundational video models have stakeholders that don’t want to be associated with what people really want to generate But they are pushing the space forward and the uncensored and unrestricted video model is coming
- giancarlostoro 1y agoNobody wants to make a commercial NSFW model that then suffers a jailbreak... for what is the most illegal NSFW content.
- yieldcrv 1y agoThats the thing, what’s “illegal” will challenge our whole society when it comes do dynamically generated real interactive avatars that are new humans When it comes to sexually explicit content in general with adults, all of our laws rely on the human actor existing FOSTA and SESTA is related to user generated content of humans, for example. They rely on making sure an actual human isnt being exploited and burdening everyone with that enforcement. When everyone can just say “thats AI” nobody’s going to care and platforms will be willing to take that risk of it being true again - or a new hit platform will. That kind of content currently Doesnt exist in large quantities yet, until a video model ungimped can generate it. Concerns about trafficking only rely on actual humans not entirely new avatars regarding children there are more restrictions that may already cover this, there is a large market for just adult looking characters though and worries about underage can be tackled independently. or be found entirely futile. not my problem, focus on what you can control. this is whats coming though. people already dont mind parasocial relationships with generative AI and already pay for that, just add nudity
- simianwords 1y agoWhy is this illegal btw? I mean whats stopping an AI company from releasing a proper NSFW model? I hope it doesn't happen but I want to know what prevents them from doing it now.
- ACCount37 1y agoThe raw model scale is not increasing by much lately. AI companies are constrained by what fits in this generation of hardware, and waiting for the next generation to become available. Models that are much larger than the current frontier are still too expensive to train, and far too expensive to serve them en masse. In the meanwhile, "better data", "better training methods" and "more training compute" are the main ways you can squeeze out more performance juice without increasing the scale. And there are obvious gains to be had there.
- xnx 1y ago> AI companies are constrained by what fits in this generation of hardware, and waiting for the next generation to become available. Does this apply to Google that is using custom built TPUs while everyone else uses stock Nvidia?
- ACCount37 1y agoBy all accounts, what's in Google's racks right now (TPU v5e, v6e) is vaguely H100-adjacent, in both raw performance and supported model size. If Google wants anything better than that? They, too, have to wait for the new hardware to arrive. Chips have a lead time - they may be your own designs, but you can't just wish them into existence.
- xxpor 1y agoAren't chips + memory constrained by process + reticle size? And therefore, how much HBM you can stuff around the compute chip? I'd expect everyone to more or less support the same model size at the same time because of this, without a very fundamentally different architecture.
- robwwilliams 1y agoThe jump to 1 million token length context for Sonnet 4 plus access to internet has been a game-changer for me. And somebody should remind Anthropic leadership to at least mirror Wikipedia; better yet support Wikipedia actively. All of the big AI players have profited from Wikipedia, but have they given anything back, or are they just parasites on FOSS and free data?
- gmadsen 1y agoIts not clear to me that it needs to. If at the margins it can still provide an advantage in the market or national defense, then the spice must flow
- duxup 1y agoI suspect it needs to if it is going to cover the costs of training.
- derefr 1y ago> Anecdotally moving from model to model I'm not seeing huge changes in many use cases. Probably because you're doing things that are hitting mostly the "well-established" behaviors of these models — the ones that have been stable for at least a full model-generation now, that the AI bigcorps are currently happy keeping stable (since they achieved 100% on some previous benchmark for those behaviors, and changing them now would be a regression per those benchmarks.) Meanwhile, the AI bigcorps are focusing on extending these models' capabilities at the edge/frontier, to get them to do things they can't currently do. (Mostly this is inside-baseball stuff to "make the model better as a tool for enhancing the model": ever-better domain-specific analysis capabilities, to "logic out" whether training data belongs in the training corpus for some fine-tune; and domain-specific synthesis capabilities, to procedurally generate unbounded amounts of useful fine-tuning corpus for specific tasks, ala AlphaZero playing unbounded amounts of Go games against itself to learn on.) This means that the models are getting constantly bigger. And this is unsustainable. So, obviously, the goal here is to go through this as a transitionary bootstrap phase, to reach some goal that allows the size of the models to be reduced. IMHO these models will mostly stay stable-looking for their established consumer-facing use-cases, while slowly expanding TAM "in the background" into new domain-specific use-cases (e.g. constructing novel math proofs in iterative cooperation with a prover) — until eventually, the sum of those added domain-specific capabilities will turn out to have all along doubled as a toolkit these companies were slowly building to "use models to analyze models" — allowing the AI bigcorps to apply models to the task of optimizing models down to something that run with positive-margin OpEx on whatever hardware that would be available at that time 5+ years down the line. And then we'll see them turn to genuinely improving the model behavior for consumer use-cases again; because only at that point will they genuinely be making money by scaling consumer usage — rather than treating consumer usage purely as a marketing loss-leader paid for by the professional usage + ongoing capital investment that that consumer usage inspires.
- kdmtctl 1y agoYou have just described a singularity point for this line of business. Which could happen. Or not.
- 1y ago
- darepublic 1y agoI hope you're right.
- wslh 1y ago> Anecdotally moving from model to model I'm not seeing huge changes in many use cases. I can just pick an older model and often I can't tell the difference... Model specialization. For example a model with legal knowledge based on [private] sources not used until now.
- dvfjsdhgfv 1y ago> I can just pick an older model and often I can't tell the difference... Or, as in the case of a leading North American LLM provider, I would love to be able to choose an older model but it chooses it for me instead.
- paulddraper 1y agoReductive. Doesn’t explain Deepseek.
- FergusArgyll 1y agoDeepseek story was way overblown. Read the gpt-oss paper, the actual training run is not the only expense. You have multiple experimental training runs as well as failed training runs. + they were behind SOTA even then
- scellus 1y agoSo far it doesn't seem like winner-take-all, and all the major players (OpenAI, Anthropic, xAI, Google, Meta?) are backed by strong partnerships and a lot of capital. It is capital-intensive this round though, so the primary producers are big and few. As long as they compete, benefits mostly go to other parties (= society) through increased productivity.
- nradov 1y agoThat's why wealthy investors connected to the AI industry are also throwing a lot of money into power generation startups, particularly fusion power. I doubt that any of them will actually deliver commercially viable fusion reactors but hope springs eternal.
- mapt 1y agoContinuing to carve out economies of scale in battery + photovoltaic for another ten doublings has plenty of positive externalities. The problem is that in the meantime, they're going to nuke our existing powergrid, created in the 1920's to 1950's to serve our population as it was in the 1970's, and for the most part not expanded since. All of the delta is in price-mediated "demand reduction" of existing users.
- UltraSane 1y agoA lot of the biggest data centers being built are also building behind the meter generation dedicated to them.
- Workaccount2 1y agoWhich is mostly natural gas sadly.
- UltraSane 1y agoYep they are tapping directly into main pipelines.
- vrt_ 1y agoImagine solving energy as a side effect of this compute race. There's finally a reason for big money to be invested into energy infrastructure and innovation to solve a problem that can't be solved with traditional approaches.
- bobsmooth 1y ago
- throw310822 1y agoJust in case, can they be repurposed for bitcoin mining? :) Edit: for the curious, no. An H100 costs about ~25k and produces $1.2/day mining bitcoin. Without factoring in electricity.
- me551ah 1y agoAnd distillation makes the compute moat irrelevant. You could spend trillions to train a model, but some companies is going to get enough data from your model and distill it's own at a much cheaper upfront cost. This would allow them to offer them for cheaper inference cost too, totally defeating the point of spending crazy money on training.
- fredoliveira 1y agoA couple of counter-arguments: Labs can just step up the way they track signs of prompts meant for model distillation. Distillation requires a fairly large number of prompt/response tuples, and I am quite certain that all of the main labs have the capability to detect and impede that type of use if they put their backs into it. Distillation doesn't make the compute moat irrelevant. You can get good results from distillation, but (intuitively, maybe I'm wrong here because I haven't done evals on this myself) you can't beat the upstream model in performance. That means that most (albeit obviously not all) customers will simply gravitate toward the better performing model if the cost/token ratio is aligned for them. Are there always going to be smaller labs? Sure, yes. Is the compute mote real, and does it matter? Absolutely.
- serf 1y ago>Labs can just step up the way they track signs of prompts meant for model distillation. Distillation requires a fairly large number of prompt/response tuples, and I am quite certain that all of the main labs have the capability to detect and impede that type of use if they put their backs into it. ....while degrading their service for paying customers. This is the same problem as law-enforcement-agency forwarding threats and training LLMs to avoid user-harm -- it's great if it works as intended, but more often than not it throws a lot more prompt cancellations at actual users by mistake, refuses queries erroneously -- and just ruins user experience. i'm not convinced any of the groups can avoid distillation without ruining customer experience.
- scottLobster 1y agoRoughly 1% of US GDP in 2025 was data center construction, mostly for AI.
- jayd16 1y agoIn this imaginary timeline where initial investments keep increasing this way, how long before we see a leak shutter a company? Once the model is out, no one would pay for it, right?
- marcosdumay 1y agoIn this imaginary reality where LLMs just keep getting better and better, all that a leak means is that you will eat-up your capital until you release your next generation. And you will want to release it very quickly either way, and should have a problem for a few months at most. And if LLMs don't keep getting qualitatively more capable every few months, that means that all this investment won't pay off and people will soon just use some open weights for everything.
- jsheard 1y agoWhatever happens if/when a flagship model leaks, the legal fallout would be very funny to watch. Lawyers desperately trying to thread the needle such that training on libgen is fair use, but training on leaked weights warrants the death penalty.
- wmf 1y agoYou can't run Claude on your PC; you need servers. Companies that have that kind of hardware are not going to touch a pirated model. And the next model will be out in a few months anyway.
- DebtDeflation 1y agoThe wildest part is that the frontier models have a lifespan of 6 months or so. I don't see how it's sustainable to keep throwing this kind of money at training new models that will be obsolete in the blink of an eye. Unless you believe that AGI is truly just a few model generations away and once achieved it's game over for everyone but the winner. I don't.
- jononor 1y agoIt is being played like a winner-takes-it-all right now (it may or may not be such a market). So it is a game of being the one that is left standing, once the others fall off. In this kind of game, speeding more is done as a strategy to increase the chances of other competitors running out of cash or otherwise hitting a wall. Sustainability is the opposite of the goal being pursued... Whether one reaches "AGI" is not considered important either, as long as one can starve out most competitors. And for the newcomers, the scale needs to be bigger than what the incumbents (Google and Microsoft) have as discretionary spending - which is at least a few billion per year. Because at that rate, those companies can sustain it forever and would be default winners. So I think yearly expenditure is going to be 20B year++
- leptons 1y agoIt's the Uber business plan - losing money until the competition loses more and goes out of business. So far Lyft seems to be doing okay, which proves the business plan doesn't really work.
- Workaccount2 1y agoThere are endless examples of that business model working...
- oblio 1y agoAre there? Which ones? I'm especially interested in companies that weren't built to be sold.
- worldsayshi 1y agoAnd we're still sort of on the fence if it's even that useful? Like sure it saves me a bit of time here and there but will scaling up really solve the reliability issues that is the real bottleneck.
- bravetraveler 1y agoAssuming the best case: we're going to need to turn this productivity into houses or lifestyle improvement, soon... or I'm just going out with Sasquatch
- worldsayshi 1y agoWhile decoding your comment I'm going to assume Sasquatch to be a semi-underground (no web site, only calls) un-startup that specializes in survival kits for people leaving civilization behind. Like calling the vacuum repair store but more hippie themed.
- bravetraveler 1y agoThat'll do :) edit: I assure you, there will still be a van
- worldsayshi 1y agoSolar powered e-van? I found this now: https://soleva.org https://soleva.org
- SchemaLoad 1y agoI feel like it's pretty settled that they are a little bit useful, as a faster search engine, or being able to automatically sort my emails. But the value is nowhere near justifying the investment.
- worldsayshi 1y agoYeah, I think they are both highly over hyped and quite undervalued in terms of how they can be used effectively.
- docdeek 1y ago> The compute moat is getting absolutely insane. We're basically at the point where you need a small country's GDP just to stay in the game for one more generation of models. For what it is worth, $13 billion is about the GDP of Somalia (about 150th in nomimal GDP) with a population of 15 million people.
- Aeolun 1y agoAs a fun comparison, because I saw the population is more or less the same. The GDP of the Netherlands is about $1.2 trillion with a population of 18 million people. I understand that that’s not quite what’s meant with ‘small country’ but in both population and size it doesn’t necessarily seem accurate.
- Aurornis 1y agoCountry scale is weird because it has such a large range. California (where Anthropic is headquartered) has over twice as many people as all of Somalia. The state of California has a GDP of $4.1 Trillion. $13 billion is a rounding error at that scale. Even the San Francisco Bay Area alone has around half as many people as Somalia.
- lofaszvanitt 1y agoNvidia needs to grow.
- asveikau 1y agoThis sounds terrible for the environment.
- 2OEH8eoCRo0 1y agoA lot of moats are just money. Money to buy competition, capture regulation, buy exclusivity, etc.
- willvarfar 1y agoAs humans don't actually work like LLMs do, we can surmise that there are far more efficient ways to get to AGI. We just need to find them.
- ijidak 1y agoCan you elaborate? The technology to build a human brain would cost billions in today’s dollars. Are you thinking moreso about energy efficiency?
- robotresearcher 1y agoWe make hundreds of millions of brains a year for the cost of their parent’s food and shelter. That’s the known minimum cost. We have a lot of room to get costs down if we can figure out how.
- xnx 1y ago> The technology to build a human brain would cost billions in today’s dollars I'm reminded of how insanely complex the human brain is: ~100 trillion connections. The Nvidia H100 has just 0.08 trillion transistors.
- senko 1y ago> We're basically at the point where you need a small country's GDP just to stay in the game for one more generation of models. When you consider where most of that money ends up (Jensen &co), it's bizarre nobody can really challenge their monopoly - still.
- xbmcuser 1y agoThis is why I keep harping on the world needing China to get competitive on node size and crashing the market. They are already making energy with solar and renewable practically free. So the world needs AI to get out of the hand of the rich few and into the hands of everyone
- derefr 1y ago> privatization You think any of these clusters large enough to be interesting, aren't authorized under a contractual obligation to run any/all submitted state military/intelligence workloads alongside their commercial workloads? And perhaps even to prioritize those state-submitted workloads, when tagged with flash priority, to the point of evicting their own workloads? (This is, after all, the main reason that the US "Framework for Artificial Intelligence Diffusion" was created: America believed China would steal time on any private Chinese GPU cluster for Chinese military/intelligence purposes. Why would they believe that? Probably because it's what the US thought any reasonable actor would do, because it's what they were doing.) These clusters might make private profits for private shareholders... but so do defense subcontractors.
- madduci 1y agoAnd just now came the email with the changes to their terms of usage and policy. Nice timing? I am sure they have scored a deal with the selling of personal data
- maqp 1y ago>You can have all the talent in the world but if you can't get 100k H100s and a dedicated power plant, you're out. I really have to wonder, how long will it be before the competition moves into who has the most wafer-scale engines. I mean, surely the GPU is a more inefficient packaging form factor than large dies with on-board HBM, with a massive single block cooler?
- mfro 1y agoSentiment I have heard is manufactories do not want to increase die size because defects per die increases at the same time.
- Workaccount2 1y agoMeanwhile at Cerebras...heh But I do believe that their cost per compute is still far more than disparate chips.
- 15155 1y agoThis is why chiplets are used.
- risyachka 1y ago>> The compute moat is getting absolutely insane. how so? deepseek and others do models on par with previous generation for a tiny fraction of a cost. Where is the moat?
- AlienRobot 1y agoI saw a story posted on reddit that U.S. engineers went to China and said the U.S. would lose the A.I. game because THE ENERGY GRID was much worse than China's. That's just pure insanity to me. It's not even Internet speed or hardware. It's literally not having enough electricity. What is going on with the world...
- ipython 1y agoNot to mention water for cooling. Large data centers can use 1 million+ gallons per day.
- xnx 1y ago1 million gallons is approximately 0.5 seconds of flow of the Columbia river.
- wiredpancake 1y agoIt means nothing when most water is recycled anyways. It's not like the GPUs actually drink the stuff, the water just connects to heatsinks and is cycled around.
- xnx 1y agoThat's true for closed loop systems, but some data enters use evaporative cooling because it is more energy efficient.
- sidewndr46 1y agoI'm not an expert at how private investment rounds work, but aren't most "raises" of AI companies just huge commitments of compute capacity? Either pre-existing or build-out.
- serf 1y agoit's difficult for me to imagine this level of compute existing and sitting there idle somewhere; it just doesn't make sense. So we can at least assume that whoever is deciding to move the capacity does so at some business risk elsewhere.
- sidewndr46 1y agoI've stayed away from the hyperscalers but worked at places where requesting 400 servers for a task was normal and routine. Understanding scale is a weird thing, that I guess is psychological. I think the different experiences and travels I have made have an impact on this. Despite living for over a decade in Texas I live in one of the most densely populated places. But I recently got to visit Colorado, which is far less populated and has lots of weird places. You can drive up to the base of the Great Sand Dunes and walk up the first few hills quite easily if you're in good shape. Here's some photos https://www.hydrogen18.com/p/2024-great-sand-dunes-national-park.html https://www.hydrogen18.com/p/2024-great-sand-dunes-national-... If you pull out your smartphone and look at Google Maps, it becomes pretty obvious how insane the scale of the place is. There's no real prohibition on where you can walk there because it isn't necessary. It's so large and the environment is so harsh, you aren't going to cross much of it on foot. Ever.
- Razengan 1y agoBarely 50 years ago computers used to cost a million dollars and were less powerful than your phone's SIM card. > GPT-4 training was what, $100M? GPT-5/Opus-4 class probably $1B+? Your brain? Basically free *(not counting time + food) Disruption in this space will come from whomever can replicate analog neurons in a better way. Maybe one day you'll be able to Matrix information directly into your brain and know kung-fu in an instant. Maybe we'll even have a Mentat social class.
- jcranmer 1y ago> Barely 50 years ago computers used to cost a million dollars and were less powerful than your phone's SIM card. Fifty years ago, we were starting to see the very beginning of workstations (not quite the personal computer of modern days), something like this: https://en.wikipedia.org/wiki/Xerox_Alto https://en.wikipedia.org/wiki/Xerox_Alto, which cost ~$100k in inflation-adjusted money.
- psychoslave 1y agoYeah, no hate for kung fu here, but maybe learning to better communicate together, act in ways that allows everyone to thrive in harmony and spread peace among all humanity might be a better thing to start incorporating, might not it?
- Razengan 1y agoIt's literally a scene from The Matrix.
- psychoslave 1y agoYes it is. We can also maybe agree that the comment wasn't implying otherwise? I mean, it's like the djin giving you three whishes, and not a single character will ask "what's the two best wishes I can do to (ensure mankind will reach perpetually best peaceful harmonious flourishing social dynamics forever| whatever goal the character might have as greatest hope)". When you have a instant perfect knowledge acquisition machine at disposal, the first thing to obviously understand is what the most important things to do to reach your goal. The film didn't mention everything Neo learned like that though, just that he accumulate straight forward for many hours. Wouldn't be an action movie, certainly you would hope the character first words after such an impressive feat wouldn't be "I know kung fu".
- huevosabio 1y agoInstead of enriching uranium we're enriching weights!
- puchatek 1y agoAnd how much will one query cost you once the companies start to try and make this stuff profitable?
- ericmcer 1y agoCould they vastly reduce this cost by specializing models? Like is a general know everything model exponentially more expensive than one that deeply understands a single topic (like programming, construction, astrophysics, whatever)? Is there room for a smaller team to beat Anthropic/OpenAI/etc. at a single subject matter?
- rich_sasha 1y agoIt's the SV playbook: invent a field, make it indispensable, monopolise it and profit. It still amazes me that Uber, a taxi company, is worth however many billions. I guess for the bet to work out, it kinda needs to end in AGI for the costs to be worth it. LLMs are amazing but I'm not sure they justify the astronomical training capex, other than as a stepping stone.
- lotsofpulp 1y agoWhy would a global taxi/delivery broker not be worth billions? Their most recent 10-Q says they broker 36 million rides or deliveries per day. Even profiting $1 on each of those would result in a company worth billions.
- simianwords 1y agoSV playbook has been to make sustainable businesses. Uber makes profits, so do Google, Amazon and other big tech. > LLMs are amazing but I'm not sure they justify the astronomical training capex, other than as a stepping stone. They can just... stop training today and quickly recuperate the costs because inference is mostly profitable.
- rich_sasha 1y agoAll these businesses looked incredibly unsustainable for a long time. Uber was a cash shredder. Amazon didn't turn a profit for years, IIRC. They became profitable essentially by becoming quasi-monopolies. Indeed, LLM companies likely turn operating profits, but I'm not sure that alone justifies their valuations. It's one thing to make money, it's another to make a return for investors. And sure, valuations are growing faster than you can blink. Time will show if this in turn is justifiable or a bubble.
- filoleg 1y agoCannot speak for the rest, but the whole “Amazon didn’t turn a profit for years” (as an argument about their profitability now coming solely through quasi-monololy routes) is incredibly misleading and bordering on disingenuous. Since before AWS was even a thing, Amazon was already turning up great revenue and could’ve easily just stopped expanding and investing into the company growth, and they would be profitable easily. Instead, Amazon decided to reinvest all their potential profits into growth/expansion (with the favorable tax treatment on top) at the expense of keeping the cash profits. At any given point, Amazon could’ve stopped reinvesting all potential profits into their growth, and they would be instantly profitable. This is not the same as Uber, which ran their core service operations at a net loss (and was only cheap due to their investors eating the difference and hoping that Uber will eventually figure out how to not lose money on operating their core service).
- matthewdgreen 1y agoWhat’s the hardware capability doubling rate for GPUs in clusters? Or (since I know that’s complicated to answer for dozens of reasons): on average how many months has it been taking for the hardware cost of training the previous generation of models to halve, excluding algorithmic improvements?
- AlexandrB 1y agoThe whole LLM era is horrible. All the innovation is coming "top-down" from very well funded companies - many of them tech incumbents, so you know the monetization is going to be awful. Since the models are expensive to run it's all subscription priced and has to run in the cloud where the user has no control. The hype is insane, and so usage is being pushed by C-suite folks who have no idea whether it's actually benefiting someone "on the ground" and decisions around which AI to use are often being made on the basis of existing vendor relationships. Basically it's the culmination of all the worst tech trends of the last 10 years.
- simianwords 1y agoThis is very pessimistic take. Where else do you think the innovation would come from? Take cloud for example - where did the innovation come from? It was from the top. I have no idea how you came to the conclusion that this implies monetization is going to be awful. How do you know models are expensive to run? They have gone down in price repeatedly in the last 2 years. Why do you assume it has to run in the cloud when open source models can perform well? > The hype is insane, and so usage is being pushed by C-suite folks who have no idea whether it's actually benefiting someone "on the ground" and decisions around which AI to use are often being made on the basis of existing vendor relationships There are hundreds of millions of chatgpt users weekly. They didn't need a C suite to push the usage.
- AlexandrB 1y ago> I have no idea how you came to the conclusion that this implies monetization is going to be awful. Because cloud monetization was awful. It's either endless subscription pricing or ads (or both). Cloud is a terrible counter-example because it started many awful trends that strip consumer rights. For example "forever" plans that get yoinked when the vendor decides they don't like their old business model and want to charge more.
- simianwords 1y agoVast majority of cloud users use AWS, GCP and Azure which have metered billing. I'm not sure what you are talking about.
- sjapkee 1y agoThe biggest problem is that result doesn't worth spent resources
- SilverElfin 1y agoThe other problem is that big companies can take a loss and starve out any competition. They already make a ton of money from various monopolies. And they do not have the distraction of needing to find funding continuously. They can just keep selling these services at a loss until they’re the only ones left. That’s leaving aside the advantages they have elsewhere - like all the data only they can access for training. For example, it is unfair that Google can use YouTube data, but no one else can. How can that be fair competition? And they can also survive copyright lawsuits with their money. And so on.
- ants_everywhere 1y ago> What gets me is that this isn't even a software moat anymore - it's literally just whoever can get their hands on enough GPUs and power infrastructure. I'm curious to hear from experts how much this is true if interpreted literally. I definitely see that having hardware is a necessary condition. But is it also a sufficient condition these days? ... as in is there currently no measurable advantage to having in-house AI training and research expertise? Not to say that OP meant it literally. It's just a good segue to a question I've been wondering about.
- powerapple 1y agoAlso not all compute was necessary for the final model, a large chunk of it is trial and error research. In theory, for $1B you spent training the latest model, a competitor will be able to do it after 6 months with $100M.
- SchemaLoad 1y agoNot only are the actual models rapidly devaluing, the hardware is too. Spend $1B on GPUs and next year there's a much better model out that's massively devalued your existing datacenter. These companies are building mountains of quicksand that they have to constantly pour more cash on else they be reduced to having no advantage rapidly.
- utyop22 1y agoYes indeed if we look at it from this equation: FCFF = EBIT(1-t) - Reinvestment If the hardware needs constant replacement, that Reinvestment number will always remain higher than what most people think. In fact, it seems none of these investments are fixed. Therefore there are no economies of scale (as it stands right now).
- chermi 1y agoIgnoring energy costs(!), I'm interested in the following. Say every server generation from nvda is 25% "better at training", by whatever metric (1). Could you not theoretically wire together 1.25 + delta more of the previous generation to get the same compute? The delta accounts for latency/bandwidth from interconnects. I'm guessing delta is fairly large given my impression of how important HBM and networking are. I don't know the efficiency gains per generation, but let's just say to get the same compute with this 1.25+delta system requires 2x energy. My impression is that while energy is a substantial cost, the total cost for a training run is still dominated by the actual hardware+infrastructure. It seems like there must be some break even point where you could use older generation servers and come out ahead. Probably everyone has this figured out and consequently the resale value of previous gen chips is quite high? What's the lifespan at full load of these servers? I think I read coreweave deprecates them (somewhat controversially) over 4 years. Assuming the chips last long enough, even if they're not usable for LLM training/serving inference, can't they be reused for scientific loads? I'm not exactly old, but back in my PhD days people were building our own little GPU clusters for MD simulations. I don't think long MD simulations are the best use of compute these days, but there's many similar problems like weather modeling, high dimensional optimization problems, materials/radiation studies, and generic simulations like FEA or simply large systems of ODEs. Are these big clusters being turned into hand-me-downs for other scientific/engineering problems like above, or do they simply burn them out? What's a realistic expected lifespan for a B200? Or maybe it's as simple as they immediately turn their last gen servers over to serve inference? Lot of questions, but my main question is just how much the hardware is devalued once it becomes previous gen. Any guidance/references appreciated! Also, anyone still in the academic computing world, do people like de shaw still exist trying to run massive MD simulations or similar? Do the big national computing centers use the latest greatest big Nvidia AI servers or something a little more modest? Or maybe even they're still just massive CPU servers? While I have anyone who might know, whatever happened to that fad from 10+ years ago saying a lot of compute/algorithms would be shifting toward more memory-heavy models(2). Seems like it kind of happened in AI at least. (1) Yes I know it's complicated, especially with memory stuff. (2) I wanna say it was ibm Almaden championing the idea.
- belter 1y agoThe AI story is over. One more unimpressive release of ChatGPT or Claude, another 2 Billion spent by Zuckerberg on subpar AI offers, and the final realization by CNBC that all of AI right now...Is just code generators, will do it. You will have ghost data centers in excess like you have ghost cities in China.
- itronitron 1y agoHmm, I wonder how much bitcoin someone could mine with that amount of compute.
- wiredpancake 1y agoA lot, but maybe a lot less than you expect. You'd be competing with ASIC miners, which are 100x more cost effective per MH/s. You don't need 100,000GB of VRAM when mining GPU, therefore its waste.
- andrewgleave 1y ago> “There's kind of like two different ways you could describe what's happening in the model business right now. So, let's say in 2023, you train a model that costs 100 million dollars. > > And then you deploy it in 2024, and it makes $200 million of revenue. Meanwhile, because of the scaling laws, in 2024, you also train a model that costs a billion dollars. And then in 2025, you get $2 billion of revenue from that $1 billion, and you spend $10 billion to train the model. > > So, if you look in a conventional way at the profit and loss of the company, you've lost $100 million the first year, you've lost $800 million the second year, and you've lost $8 billion in the third year. So, it looks like it's getting worse and worse. If you consider each model to be a company, the model that was trained in 2023 was profitable.” > ... > > “So, if every model was a company, the model is actually, in this example, is actually profitable. What's going on is that at the same time as you're reaping the benefits from one company, you're founding another company that's like much more expensive and requires much more upfront R&D investment. And so, the way that it's going to shake out is this will keep going up until the numbers go very large, the models can't get larger, and then it will be a large, very profitable business, or at some point, the models will stop getting better. > > The march to AGI will be halted for some reason, and then perhaps it will be some overhang, so there will be a one-time, oh man, we spent a lot of money and we didn't get anything for it, and then the business returns to whatever scale it was at.” > ... > > “The only relevant questions are, at how large a scale do we reach equilibrium, and is there ever an overshoot?” From Dario’s interview on Cheeky Pint: https://podcasts.apple.com/gb/podcast/cheeky-pint/id1821055332?i=1000720897619 https://podcasts.apple.com/gb/podcast/cheeky-pint/id18210553...
- protocolture 1y ago>The compute moat is getting absolutely insane. Is it? Seems like theres a tiny performance gain between "This runs fine on my laptop" and "This required a 10B dollar data centre" I dont see any moat, just crazy investment hoping to crack the next thing and moat that.
- cjbgkagh 1y agoThat’s like being upset that you can’t dig your own suez canal. So long as there is competition it’ll be available at marginal cost. And there is plenty of innovation that can be done on the edges, and not all of machine learning is LLMs.
- mlyle 1y ago> So long as there is competition it’ll be available at marginal cost. Most things are not perfect competition, so you get MR=MC not P=MC. We're talking about massive capital costs. Another name for massive capital costs are "barriers to entry".
- cjbgkagh 1y agoGranted that capital costs are a barrier to entry and that barriers to entry leads to non-perfect competition, but the exploitability is limited in the case of LLMs because they exist on a sub-linear utility scale. In LLMs 2x the price is not 2x as useful, this means a new entrant can enter the lower end of the market and work their way up. The only way to prevent that is for the incumbent to keep costs as close to marginal as possible. There is a natural monopoly aspect given the ability to train and data mine on private usage data but in general improvements in the algorithms and training seem to be dominating advancements. Microsoft's search engine Bing paid an absolute fortune for access to usage data and they were unable to capitalize on it. LLMs have the unusual property that a lot of value can be extracted out of fine tuning for a specialized purposes which opens the door to a million little niches providing fertile ground for future competitors. This is one area where being a fast follower makes a lot of sense.
- mlyle 1y agoAlmost anything has a utility scale which is diminishing. But we still see MR=MC pricing in industries with barriers to entry (IPR, capital costs). TSMC and Mercedes don't price cheap to avoid giving others a toehold. > There is a natural monopoly aspect given the ability to train and data mine on private usage data but in general improvements in the algorithms and training seem to be dominating advancements. There's pretty big economies of scale with inference-- the magic of how to route correctly with experts to conduct batching while keeping latency low. It's an expensive technology to create, and there's a large minimum scale where it works well.
- noosphr 1y agoMy hope is that this hype cycle overbuilds nuclear power capacity so much that we end up using it to sequester carbon dioxide from the atmosphere once the bubble pops and electricity prices become negative for most of the day. In the medium term China has so much spare capacity that they maybe be the only game in town for highend models, while the US will be trying to fix a grid with 50 years of deferred maintenance.
- lz400 1y agoThat’s probably what the companies spending the money think, that they’re building a huge moat. There’s an alternative view. If there’s a bubble and all these companies are spending these huge sums on something that ends up not returning that much on that investment, and the models plateau and eventually smaller, cheaper, self-runnable open source versions get 90% of the way there, what’s going to happen to that moat? And the companies that over spent so much? This article is a good example of the bear case https://www.honest-broker.com/p/is-the-bubble-bursting https://www.honest-broker.com/p/is-the-bubble-bursting
- BobbyTables2 1y agoUntil one day an outsider finds a new approach for LLMs that vastly reduces the computational complexity. And then we’ll realize we wasted an entire Apollo space program to build an over-complicated autocompleter.
- ath3nd 1y ago> GPT-7 will need its own sovereign wealth fund If the diminishing returns that we see now continue to prove true, ChatGPT6 will already be financially not viable so I doubt there will be GPT7 that can live up to the big version bump. Many folks already consider GPT5 to be more like GTP4.1. I personally am very bearish on Anthropic and OpenAI.
- tootie 1y agoThis is why Nvidia is the most valuable company in the world. Ultimately all these investment rounds for LLM companies are just going to be spent on Nvidia products.
- mikewarot 1y agoMost of that power usage is moving data and weights into multiply accumulate hardware, then moving the data out. The actual computation is a fairly small fraction of the power consumed. It's quite likely that an order of magnitude improvement can be had. This is an enormous incentive signal for someone to follow.
- delusional 1y ago> The compute moat Does this really describe a "most" ør are you just describing capital? The capitalization is getting insane. Were basically at the point where you ned more capital than a small nations GDP. That sounds mich more accurate to my ears, and much more troubling
- up2isomorphism 1y agoThere is no generational differences between these models. I tested cursors with all different backends and they are similar in most cases. So called race is just a Wall Street sensation to bump the stock price.
- illiac786 1y agoI sincerely hope this whole LLM monetization scheme crashes and burns down on these companies. I really hope we can get to a point where modest hardware will achieve similar results for most tasks and these insane amount of hardware will only be required for the most complex requests only, which will be rarer, thereby killing the business case. I would dance the Schadenfreude Opus in C major if that became the case.