9 ms·
Why wouldn't you factor in training? It is not like you can train once and then have the model run for years. You need to constantly improve to keep up with the
by BlindEyeHalo 1y ago
Why wouldn't you factor in training? It is not like you can train once and then have the model run for years. You need to constantly improve to keep up with the competition. The lifespan of a model is just a few months at this point.
- deleted 1y ago[deleted]
- MontyCarloHall 1y agoAs long as models continue on their current rapid improvement trajectory, retraining from scratch will be necessary to keep up with the competition. As you said, that's such a huge amount of continual CapEx that it's somewhat meaningless to consider AI companies' financial viability strictly in terms of inference costs, especially because more capable models will likely be much more expensive to train. But at some point, model improvement will saturate (perhaps it already has). At that point, model architecture could be frozen, and the only purpose of additional training would be to bake new knowledge into existing models. It's unclear if this would require retraining the model from scratch, or simply fine-tuning existing pre-trained weights on a new training corpus. If the former, AI companies are dead in the water, barring a breakthrough in dramatically reducing training costs. If the latter, assuming the cost of fine-tuning is a fraction of the cost of training from scratch, the low cost of inference does indeed make a bullish case for these companies.
- mgh95 1y ago> If the latter, assuming the cost of fine-tuning is a fraction of the cost of training from scratch, the low cost of inference does indeed make a bullish case for these companies. On the other hand, this may also turn into cost effective methods such as model distillation and spot training of large companies (similarly to Deepseek). This would erode the comparative advantage of Anthropic and OpenAI, and result in a pure value-add play for integration with data sources and features such as SSO. It isn't clear to me that a slowing of retraining will result in advantages to incumbents if model quality cannot be readily distinguished by end-users.
- echelon 1y ago> model distillation I like to think this is the end of software moats. You can simply call a foundation model company's API enough times and distill their model. It's like downloading a car. Distribution still matters, of course.
- vonneumannstan 1y agoI suspect we've already reached the point with models at the GPT5 tier where the average person will no longer recognize improvements and this model can be slightly improved at slow intervals and indeed run for years. Meanwhile research grade models will still need to be trained at massive cost to improve performance on relatively short time scales.
- AJ007 1y agoWhenever someone has complained to me about issues they are having with ChatGPT on a particular question or type of question, the first thing I do is ask them what model they are using. So far, no one has ever known offhand what model they were using, nor were not aware there are more models! If you understand there are multiple models from multiple providers, some of those models are better at certain things than others, and how you can get those models to complete your tasks, you are in the top 1% (probably less) of LLM users.
- deleted 1y ago[deleted]
- th0ma5 1y agoThis would be helpful if there was some kind of first principle at which to gauge that better or worse comparison but there isn't outside of people's value judgements like what you're offering.
- black_knight 1y agoStrangely, I feel GPT-5 as the opposite of an improvement over the previous models, and consider just using Claude for actual work. Also the voice mode went from really useful to useless “Absolutely, I will keep it brief and give it to you directly. …some wrong annswer… And there you have it! As simple as that!”
- vonneumannstan 1y ago>Strangely, I feel GPT-5 as the opposite of an improvement over the previous models This is almost surely wrong but my point was about GPT5 level models in general not GPT5 specifically...
- christina97 1y agoIn the same way that every other startup tries to sweep R&D costs under the rug and say “yeah but the marginal unit economics have 50% gross margins, we’ll be a great business soon”.
- utyop22 1y agolol. TBH I don't take anyone seriously unless they are talking about cash flows (FCFF or FCFE specifically). Who cares about expense classification - show me the money!
- danielmarkbruce 1y agoGoogle and Facebook had negative free cash flow for years early in their lives. All the good investors were lolling at the bad investors lolling at the cash they were burning.
- utyop22 1y agoOk and lets compare the cost of running those products and reinvestment vs the model businesses. FCFF = EBIT(1-t)-Reinvestment. The operating expenses of the model business are much higher - so lower EBIT. The larger the reinvestment the larger the hole. And the longer it continues (without clear steep barriers to entry to exclude competitors in the long run) it becomes harder to justify a high valuation. I really dislike comparisons like this - it glosses over a lot of details.
- danielmarkbruce 1y agoOne can explain the equation all they like - the fact is that negative free cash flow is just a reality of the early stages of some very, very good businesses. In the 90's and early 2000s, but people laughed at businesses like Amazon & Google for years. These types of people highly focused on the free cash flow of a business in it's early years are just dumb. Sometimes a business takes a lot of investment in the early stages - whether it's capex for data centers or S&M for enterprise software businesses, or R&D for pharma businesses or whatever. As for "clear steep barriers" - again, just clueless stuff. There weren't clear steep barriers to search when Google started, there were dozens of search engines. Google created them. Creating barriers to entry is expensive and the "FCFF people" imagine they arrive out of thin air. It takes a lot of time and or money to create them. It's unclear if "the model business" is going to be high or low margin. It's unclear how high the barriers to entry for making models will be in practice. It's unclear what the reinvestment required will be. We are a few years into it. About the only thing that is clear is this: if you try to run a positive free cashflow business in this space over the next few years, you'll be crushed. If you want a shot at a large, high return on capital business come 2035, you better be willing to spend up now.
- ugh123 1y agoIt's possible they factor in training purely as an "R&D" cost and then can tax that development at a lower rate.
- _sword 1y agoI spoke with management at a couple companies that were training models, and some of them expensed the model training in-period as R&D. That's why
- jacurtis 1y agoIn a recent episode of Hard Fork podcast, the hosts discussed an on-the-record conversation they had with Sam Altman from OpenAI. They asked him about profitability and he claimed that they are losing money mostly because of the cost of training. But as the model advances, they will train less and less. Once you take training out of the equation he claimed they were profitable based on the cost of serving the trained foundation models to users at current prices. Now, when he said that, his CFO corrected him and said they aren't profitable, but said "it's close". Take that with a grain of salt, but thats a conversation from one of the big AI companies that is only a few weeks old. I suspect that it is pretty accurate that pricing is currently reasonable if you ignore training. But training is very expensive and the reason most AI companies are losing money right now.
- pas 1y ago> most AI companies are losing money right now which is completely "normal" at this point, """right"""? if you have billions of VC money chasing returns there's no time to sit around, it's all in, the hype train doesn't wait for bootstrapping profitability. and of course with these gargantuan valuations and mandatory YoY growth numbers, there is no way they are not fucking with the unit economy numbers too. (biases are hard to beat, especially if there's not much conscious effort to do so.)
- brianwawok 1y agoDoes the cost of good come down 10x or not? For say Uber it didn’t, so we went from great $6 VC funded product to mediocre $24 ride product we have today. I’m not sure I’m going to use Copilot at $1 per request. Or even $0.25. Starts to approach overseas consultant in price and ability.
- pas 1y agowell, Uber always faced the obvious problem of scaling (even after level 42 self-driving, because it's not possible to serve local demand with global supply, plus all the regulatory compliance issues - which they initially "conveniently" sidestepped by being bold/criminal, but cities are not going to play dumb forever) of course these chat-AIs also started by "well maybe it's fair use", but at least the scaling problem seems easier than for taxi services