7 ms·
Berkeley Researchers Replicate DeepSeek R1's Core Tech for Just $30: A Small Mod
- aurareturn 2y agoThis is truly the biggest breakthrough from DeepSeek - that an LLM can teach itself to reason, no human feedback needed. That’s nuts and brings forward the idea that an AI is close to self improvement.
- EGreg 2y agoWe already knew that kind of stuff from AlphaZero vs AlphaGo AlphaGo Zero is a version of DeepMind's Go software AlphaGo. AlphaGo's team published an article in Nature in October 2017 introducing AlphaGo Zero, a version created without using data from human games, and stronger than any previous version.[1] By playing games against itself, AlphaGo Zero: surpassed the strength of AlphaGo Lee in three days by winning 100 games to 0; reached the level of AlphaGo Master in 21 days; and exceeded all previous versions in 40 days.[2] Training artificial intelligence (AI) without datasets derived from human experts has significant implications for the development of AI with superhuman skills, as expert data is "often expensive, unreliable, or simply unavailable."[3] Demis Hassabis, the co-founder and CEO of DeepMind, said that AlphaGo Zero was so powerful because it was "no longer constrained by the limits of human knowledge".[4] Furthermore, AlphaGo Zero performed better than standard deep reinforcement learning models (such as Deep Q-Network implementations[5]) due to its integration of Monte Carlo tree search. David Silver, one of the first authors of DeepMind's papers published in Nature on AlphaGo, said that it is possible to have generalized AI algorithms by removing the need to learn from humans.[6] Google later developed AlphaZero, a generalized version of AlphaGo Zero that could play chess and Shōgi in addition to Go.[7] In December 2017, AlphaZero beat the 3-day version of AlphaGo Zero by winning 60 games to 40, and with 8 hours of training it outperformed AlphaGo Lee on an Elo scale. AlphaZero also defeated a top chess program (Stockfish) and a top Shōgi program (Elmo).[8][9] Source: https://en.wikipedia.org/wiki/AlphaGo_Zero https://en.wikipedia.org/wiki/AlphaGo_Zero
- aurareturn 2y agoIndeed. But DeepSeek is the first to release a working repro for LLMs as far as I know.
- simlevesque 2y agoYeah but LLM and game AIs are very different. An AI can easily tell if it won the game but an LLM can't know by itself if what they said was useful.
- thorum 2y agoIt’s great, but it only works for problems where there is exactly one correct solution and it’s possible to automatically verify the solution - like math and programming. So far these reasoning models have not shown much transfer learning of reasoning to other domains, and are often worse at non-math/code tasks than standard models.
- datameta 2y agoNot sure if this a trivial or naive thought - is that perhaps because in non-discrete ideas there is more granularity of information encoded compared to a numerical solution? Separatelt, but on a related note - do we need analog or quantum computing to "truly" scale?
- cluckindan 2y agoTo ”truly” scale we need a system directly doing calculation on some fundamental property of matter. Quantum computing is _close_ but suffers from a lack of interpretability: how would we know if our quantum simulation of the universe would include the quantum computer doing the simulation, as well as another recursive copy of the simulation, and another, and another…
- onlyrealcuzzo 2y agoThere are infinite solutions to coding problems. You could have lots of comments and pass statements and unnecessary conditionals. I don't think it matters that there is only one correct answer. It matters that you can verify reliably if the answer is correct enough.
- JKCalhoun 2y agoSo, we have a very good left-brain model.
- vintermann 2y agoWe can probably make it work at more nebulous problems by using a big LLM to judge the quality of the answers as well. It should be easier to recognise e.g. a really good poetic translation than to make one, and as long as that's true it could benefit from internal monologue "reasoning" in those domains as well.
- segasaturn 2y agoThis makes ScaleAI obsolete right?
- ricksunny 2y agoA physics simulator's rules will be derivative of known physics. If we ask it to push the boundaries of known physics, then we can't verify it without real-world experiment. Absent that then it's basically Chegg on steroids: "Tell me, based on what we believe about how the universe works, about X." is the implicit preface to every physics question. To the extent that someone can figure out a way out of this epistemic box, I'm interested.
- cluckindan 2y agoIf the current hubbub around DeepSeen is really because they ”created their model” with like $5M when previously ”creating a model” cost $500B, it is rather obvious that ”creating the model” with just $30 implies the meanings of the three ”creating a model” expressions are highly divergent.
- simlevesque 2y agoNot really, the two first examples you named are the same meaning.
- drawnwren 2y agoThey aren’t though. $5M is the cost of a single training run. $500B includes the cost of operations, data center, a lot more failed runs because they weren’t sure that they’d were on the right path etc.
- onlyrealcuzzo 2y agoCompare to what you think a single run cost then. It was orders of magnitude more before DeepSeek.
- tomrod 2y agoSo far. And there might be orders of magnitude left to improve! Deepseek R1 is a training architecture improvement -- cool stuff! [0] [0] https://newsletter.languagemodels.co/p/the-illustrated-deepseek-r1 https://newsletter.languagemodels.co/p/the-illustrated-deeps...
- EA-3167 2y agoRight. It's like building a large model rocket and saying that you've cracked rocketry for a fraction of the cost that was required in the 1940's and 1950's. Well yes, yes you did, because all you had to do was follow the existing instructions, guidelines, and use easily available materials. You didn't go down any dead ends, didn't have to work your way from propellants like high test peroxide, dangerous hypergolics, and eventually develop solid rocket boosters. It's like making the generic of a drug someone else developed.
- highfrequency 2y agoFirst graph tells the story - below a certain model size (500m params), reinforcement learning is close to useless. Above this (task-dependent) model size threshold, reinforcement learning basically works. I suspect this is what we saw play out with math/coding reasoning models - until recently, the base models were not good enough for ~random output search to hit on a correct path with any reasonable frequency. Below this threshold of base model intelligence, the only efficient way forward was to collect plain supervised data (either through human labeled math problem solutions [1] or meticulous filtering of web text [2]. But as soon the base model (in this case Deepseek V3) breaks through and can actually solve a decent fraction of math problems, then reinforcement learning (plus other simple tricks like chain-of-thought prompting, simple ensemble voting, etc.) can easily juice the results through the following loop: 1) random search through different solution paths 2) identify the correct solution paths based on the final answer 3) train on the correct solution paths The exciting thing is that not only can RL bump up the performance of the current base model, but it can be used to generate new high-quality reasoning trace data, which was in painfully short-supply for training the initial models. This leads to a new wave of base models with better one-pass intuition, which leads to more efficient reinforcement learning search on harder problems, which leads to better training data... Note that this was basically impossible for non-LLM models in the past. You could always juice ImageNet classification performance with a simple ensemble of identically trained models, but that path didn't lead anywhere interesting because a juiced model didn't allow the creation of new synthetic data that was superior to the data it was trained on. The key difference is that LLMs not only output the solution but also output a solution path with all the intermediate steps - and these searched-and-filtered solution paths are much more valuable than the vast majority of the model's initial training data. [1] https://arxiv.org/abs/2305.20050 https://arxiv.org/abs/2305.20050 [2] https://arxiv.org/abs/2402.03300 https://arxiv.org/abs/2402.03300 and https://arxiv.org/abs/2206.14858 https://arxiv.org/abs/2206.14858
- krackers 2y agoYeah that is the hypothesis some people have floated for as to why we only "just now" discovered this simple idea: https://news.ycombinator.com/item?id=42837349 https://news.ycombinator.com/item?id=42837349
- DannyPage 2y agoUnless I missed it, it seems strange that the article wouldn’t link to the Github repo for the TinyZero model. https://github.com/Jiayi-Pan/TinyZero https://github.com/Jiayi-Pan/TinyZero
- semking 2y agoYou are right! I commented here and on their substack to give the credit to the OPs!
- fp64 2y ago$30 is “less than a dinner for two”?
- jedberg 2y agoFor a grad student. It's about 2 burritos in Berkeley. :)
- nick3443 2y agoWould it be correct to summarize that the general conceptual shift is optimizing MOEs on more specific smaller tasks? It smells like borderline overfitting to me for some reason.
- whimsicalism 2y agothe core insight is that you can train on extremely sparse deterministic reward signal and it just works that MoEs are better for a given compute budget has been known for a while.
- nick3443 2y agoThank you for the reply. Maybe I'm daft or out of the loop, I don't see the difference between this method and what I have seen previously as "fine tuning on synthetic data". They also mention in the source tweet that the quality of the base model is important, which is not really accounted for in the $30 figure.
- whimsicalism 2y ago'finetuning on synthetic data' = same LM task with BCE objective, ie. training on next word prediction on good text. 'rl on sparse rewards' = RL objective - reward is given by deterministic evaluation, reward is passed back as sparse signal. there is no 'ground truth' data to compare against.
- nick3443 2y agoAh now I see the difference. Thank you!!
- whimsicalism 2y ago'replication' requires matching benchmark performance, definitionally. more like 'demonstrates the technique generalizes' here. HN has really been inundated with blogspam recently
- ipsum2 2y agoReposting the comment from https://news.ycombinator.com/item?id=42843959 https://news.ycombinator.com/item?id=42843959: This is blogspam of https://github.com/Jiayi-Pan/TinyZero https://github.com/Jiayi-Pan/TinyZero and https://nitter.lucabased.xyz/jiayi_pirate/status/1882839370505621655 https://nitter.lucabased.xyz/jiayi_pirate/status/18828393705.... This also doesn't mention that it's for one specific domain (playing Countdown). See also https://news.ycombinator.com/item?id=42819262 https://news.ycombinator.com/item?id=42819262.
- oytis 2y agoIs it some kind of a joke?
- xigency 2y agoThis reads like an AI hallucination. I'm willing to steak-out this ground even if I'm wrong because of the glaring lack of skepticism. I don't even know when dollars became a concrete compute measure. We used to use FLOPs before we were trying to pull headlines like it were a claw machine game.
- semking 2y agoThe title is click-baity if you ask me. I didn't change it.
- snake_doc 2y ago@dang please link to either the GitHub https://github.com/Jiayi-Pan/TinyZero https://github.com/Jiayi-Pan/TinyZero or the primary source twitter thread: https://x.com/jiayi_pirate/status/1882839370505621655 https://x.com/jiayi_pirate/status/1882839370505621655
- semking 2y agoGuys I'm sorry but it appears the substack did NOT link to the original authors which is NOT acceptable! Credit GitHub: https://github.com/Jiayi-Pan/TinyZero https://github.com/Jiayi-Pan/TinyZero Source on X: https://x.com/jiayi_pirate/status/1882839370505621655 https://x.com/jiayi_pirate/status/1882839370505621655
- semking 2y agoI just left a comment on their Substack to give the credit to the OPs.
- excalibur 2y agoGood to know that our AI overlords will be built as cheaply as possible. If there's one thing I can't stand about bondage it's inefficiency.
- SubiculumCode 2y agoSell Nvidia last week? Seriously. Or is it that now we can make smaller models more powerful, and then run more of them to get more work done.
- UncleOxidant 2y ago"TinyZero is a reproduction of DeepSeek R1 Zero in countdown and multiplication tasks." Does that mean that this has very limited utility (to certain math problems)?