4 ms·
Not really, the two first examples you named are the same meaning.
by simlevesque 2y ago
Not really, the two first examples you named are the same meaning.
- drawnwren 2y agoThey aren’t though. $5M is the cost of a single training run. $500B includes the cost of operations, data center, a lot more failed runs because they weren’t sure that they’d were on the right path etc.
- onlyrealcuzzo 2y agoCompare to what you think a single run cost then. It was orders of magnitude more before DeepSeek.
- tomrod 2y agoSo far. And there might be orders of magnitude left to improve! Deepseek R1 is a training architecture improvement -- cool stuff! [0] [0] https://newsletter.languagemodels.co/p/the-illustrated-deepseek-r1 https://newsletter.languagemodels.co/p/the-illustrated-deeps...
- EA-3167 2y agoRight. It's like building a large model rocket and saying that you've cracked rocketry for a fraction of the cost that was required in the 1940's and 1950's. Well yes, yes you did, because all you had to do was follow the existing instructions, guidelines, and use easily available materials. You didn't go down any dead ends, didn't have to work your way from propellants like high test peroxide, dangerous hypergolics, and eventually develop solid rocket boosters. It's like making the generic of a drug someone else developed.
- deadbabe 2y agoSo what’s your point, everyone developing a model should be forced to spend the same as what the first movers did?
- skeaker 2y agoProbably just that it's not as impressive as it appears because it didn't innovate. Which is of course irrelevant since the innovative leap here were the optimizations that let them make their model with an order of magnitude fewer materials, regardless of whatever innovation costs OpenAI ate.
- nightpool 2y agoNo, just that what DeepSeek did is not as valuable as what the first movers did, because it did not advance the state of the art nearly as much. It's a new cheaper way to go from Base LM -> CoT "reasoning" LM. We already had CoT "reasoning" LMs, so while the new cheaper path to get to them is interesting, it's not necessarily groundbreaking either. Also, R1 only works with the "cold start" data that they distilled from o1, so it's not quite clear that it'll ever be able to exceed o1's capabilities. We already know it's much cheaper to distill new, smaller models from large already pretained and well-performing models—in fact, $5M sounds like a very expensive way to do so. So while these new techniques are probably going to have some impact, OpenAI is far from quaking in their boots
- deadbabe 2y agoUm it’s very valuable, maybe even more valuable… Companies can now have a private LLM with o1 quality without having to send data to OpenAI or pay for their API.
- dkjaudyeqooe 2y agoYou're more or less describing how all progress on anything, ever, happened. Even if you merely flipped a single bit and created AGI based on existing tech, you're still the legitimate creator of AGI.
- drawnwren 2y agoI don’t think anyone is claiming that Deepseek didn’t produce a very impressive frontier model. They’re just saying it’s not surprising that, in your analogy, flipping the single bit was cheaper than the prior work.
- deleted 2y ago[deleted]
- bilbo0s 2y agoThe first two are not the same thing. DeepSeek never claimed to have trained the base models they used. Now maybe a lot of people inferred that they did, but that's not what they claimed. Their breakthrough was more along the lines of, "given a model, we can train it to reason using a certain class of RL techniques". Which is, to my mind, more useful in any case. But yeah, if there are people out there thinking they can train base models from scratch for USD6 Million with no data, they're likely to be disagreeably surprised when they make the attempt.
- deleted 2y ago[deleted]
- supermatt 2y ago> DeepSeek never claimed to have trained the base models they used Isn't that exactly what they have done? Maybe you are confused with the distilled models? DeepSeek-R1: "DeepSeek-R1-Zero & DeepSeek-R1 are trained based on DeepSeek-V3-Base" DeepSeek-V3: "At an economical cost of only 2.664M H800 GPU hours, we complete the pre-training of DeepSeek-V3 on 14.8T tokens, producing the currently strongest open-source base model. The subsequent training stages after pre-training require only 0.1M GPU hours"
- chpatrick 2y agoAren't they based on qwen? I thought the clever bit with DeepSeek is a way to fine tune with reinforcement learning, not training a huge model from scratch.
- vineyardmike 2y agoNo… mostly. They developed their “zero” model, which (they claim) is a from-scratch model. Base models are typically not fine-tuned for particular applications (eg chat). They trained their Zero model into a Chat+Reasoning model, which is what is attracting news. They ALSO fine-tuned small Qwen models using their big models as a teacher (distillation technique).
- 2y ago