5 ms·
They aren’t though. $5M is the cost of a single training run. $500B includes the cost of operations, data center, a lot more failed runs because they weren’t su
by drawnwren 2y ago
They aren’t though. $5M is the cost of a single training run. $500B includes the cost of operations, data center, a lot more failed runs because they weren’t sure that they’d were on the right path etc.
- onlyrealcuzzo 2y agoCompare to what you think a single run cost then. It was orders of magnitude more before DeepSeek.
- tomrod 2y agoSo far. And there might be orders of magnitude left to improve! Deepseek R1 is a training architecture improvement -- cool stuff! [0] [0] https://newsletter.languagemodels.co/p/the-illustrated-deepseek-r1 https://newsletter.languagemodels.co/p/the-illustrated-deeps...
- EA-3167 2y agoRight. It's like building a large model rocket and saying that you've cracked rocketry for a fraction of the cost that was required in the 1940's and 1950's. Well yes, yes you did, because all you had to do was follow the existing instructions, guidelines, and use easily available materials. You didn't go down any dead ends, didn't have to work your way from propellants like high test peroxide, dangerous hypergolics, and eventually develop solid rocket boosters. It's like making the generic of a drug someone else developed.
- deadbabe 2y agoSo what’s your point, everyone developing a model should be forced to spend the same as what the first movers did?
- skeaker 2y agoProbably just that it's not as impressive as it appears because it didn't innovate. Which is of course irrelevant since the innovative leap here were the optimizations that let them make their model with an order of magnitude fewer materials, regardless of whatever innovation costs OpenAI ate.
- nightpool 2y agoNo, just that what DeepSeek did is not as valuable as what the first movers did, because it did not advance the state of the art nearly as much. It's a new cheaper way to go from Base LM -> CoT "reasoning" LM. We already had CoT "reasoning" LMs, so while the new cheaper path to get to them is interesting, it's not necessarily groundbreaking either. Also, R1 only works with the "cold start" data that they distilled from o1, so it's not quite clear that it'll ever be able to exceed o1's capabilities. We already know it's much cheaper to distill new, smaller models from large already pretained and well-performing models—in fact, $5M sounds like a very expensive way to do so. So while these new techniques are probably going to have some impact, OpenAI is far from quaking in their boots
- deadbabe 2y agoUm it’s very valuable, maybe even more valuable… Companies can now have a private LLM with o1 quality without having to send data to OpenAI or pay for their API.
- dkjaudyeqooe 2y agoYou're more or less describing how all progress on anything, ever, happened. Even if you merely flipped a single bit and created AGI based on existing tech, you're still the legitimate creator of AGI.
- drawnwren 2y agoI don’t think anyone is claiming that Deepseek didn’t produce a very impressive frontier model. They’re just saying it’s not surprising that, in your analogy, flipping the single bit was cheaper than the prior work.