5 ms·
This whole article is built off using DeepSeek R1, which is a huge premise that I don't think is correct. DeepSeek is much more efficient and I don't think it's
by sc68cal 1y ago
This whole article is built off using DeepSeek R1, which is a huge premise that I don't think is correct. DeepSeek is much more efficient and I don't think it's a valid way to estimate what OpenAI and Anthropic's costs are.
https://www.wheresyoured.at/deep-impact/ https://www.wheresyoured.at/deep-impact/
Basically, DeepSeek is _very_ efficient at inference, and that was the whole reason why it shook the industry when it was released.
- boroboro4 1y agoDeepSeek inference efficiency comes from two things: MoE and MLA attention. OpenAI was rumored to use MoE around GPT4 moment, I.e loooong time ago. Given Gemini efficiency with long context I would bet their attention is very efficient too. GPT OSS uses fp4, which DeepSeek doesn’t use yet btw. So no, big labs aren’t behind DeepSeek in efficiency. Not by much at least.
- GaggiX 1y agoThe "efficiency" meantioned in blog post you have linked is the price difference between Deepseek and o1, it doesn't mean that GPT-5 or other SOTA models are less efficient.
- CjHuber 1y agoThe reason it shook the market at least was because of the claim that its training cost was 5 million.
- hirako2000 1y agoThat' what the buzz focused on, strange as we don't actually know what it cost them. While inference optimization is a fact and is even more impactful since training costs benefit from economics of scale.
- CjHuber 1y agoI don't think that's strange at all, it's a much more palatable narrative for the mass who doesn't know what inference and training is and who think having conversations=training
- hirako2000 1y agoI agree nothing surprising in that, also back then inference wasn't as much questioned as today with regards to being sold at a loss.
- vitaflo 1y agoAlso the fact that it cost 10% of what other models cost. Pretty much still does.
- phillipcarter 1y agoUhhh, I'm pretty sure DeepSeek shook the industry because of a 14x reduction in training cost, not inference cost. We also don't know the per-token cost for OpenAI and Anthropic models, but I would be highly surprised if it was significantly more expensive than open models anyone can use and run themselves. It's not like they're also not investing in inference research.
- andai 1y agoIsn't training cost a function of inference cost? From what I gathered, they reduced both. I remember seeing lots of videos at the time explaining the details, but basically it came down to the kind of hardware-aware programming that used to be very common. (Although they took it to the next level by using undocumented behavior to their advantage.)
- booi 1y agoThey're typically somewhat related but the difference between training and inference can vary greatly so, i guess the answer is no. they did reduce both though and mostly due to reduced precision
- baxtr 1y agoBecause of the alleged reduction in training costs.
- basilgohar 1y agoAll reports by companies are alleged until verified by other, more trustworthy sources. I don't think it's especially notable that it's alleged because it's DeepSeek vs. the alleged numbers from other companies.
- gmd63 1y agoDeepSeek was trained with distillation. Any accurate estimate of training costs should include the training costs of the model that it was distilling.
- davidguetta 1y agoWhat a wrong take. Its not even MoE that was great in deepseek, its shared expert + grpo
- thatguysaguy 1y agoWhy would you think that deepseek is more efficient than gpt-5/Claude 4 though? There's been enough time to integrate the lessons from deepseek.
- overgard 1y agoBecause to make GPT-5 or Claude better than previous models, you need to do more reasoning which burns a lot more tokens. So, your per-token costs may drop, but you may also need a lot more tokens.
- jstummbillig 1y agoGPT-5 can be configured extensively. Is there any point at which any configuration of GPT-5 that offers ~DeepSeek level performance is more expensive than DeepSeek per token?
- dcre 1y agoWhat are we meant to take away from the 8000 word Zitron post? In any case, here is what Anthropic CEO Dario Amodei said about DeepSeek: "DeepSeek produced a model close to the performance of US models 7-10 months older, for a good deal less cost (but not anywhere near the ratios people have suggested)" "DeepSeek-V3 is not a unique breakthrough or something that fundamentally changes the economics of LLM’s; it’s an expected point on an ongoing cost reduction curve. What’s different this time is that the company that was first to demonstrate the expected cost reductions was Chinese." https://www.darioamodei.com/post/on-deepseek-and-export-controls https://www.darioamodei.com/post/on-deepseek-and-export-cont... We certainly don't have to take his word for it, but the claim is that DeepSeek's models are not much more efficient to train or inference than closed models of comparable quality. Furthermore, both Amodei and Sam Altman have recently claimed that inference is profitable: Amodei: "If you consider each model to be a company, the model that was trained in 2023 was profitable. You paid $100 million, and then it made $200 million of revenue. There's some cost to inference with the model, but let's just assume, in this cartoonish cartoon example, that even if you add those two up, you're kind of in a good state. So, if every model was a company, the model, in this example, is actually profitable. What's going on is that at the same time as you're reaping the benefits from one company, you're founding another company that's much more expensive and requires much more upfront R&D investment. And so the way that it's going to shake out is this will keep going up until the numbers go very large and the models can't get larger, and then it'll be a large, very profitable business, or, at some point, the models will stop getting better, right? The march to AGI will be halted for some reason, and then perhaps it'll be some overhang. So, there'll be a one-time, 'Oh man, we spent a lot of money and we didn't get anything for it.' And then the business returns to whatever scale it was at." https://cheekypint.substack.com/p/a-cheeky-pint-with-anthropic-ceo https://cheekypint.substack.com/p/a-cheeky-pint-with-anthrop... Altman: "If we didn’t pay for training, we’d be a very profitable company." https://www.theverge.com/command-line-newsletter/759897/sam-altman-chatgpt-openai-social-media-google-chrome-interview https://www.theverge.com/command-line-newsletter/759897/sam-...
- gmerc 1y agoGrok 3.5: 400M training run DeepSeek R1: 5M training run Released around the same time, marginal performance difference.