10 ms·
R-Zero: Self-Evolving Reasoning LLM from Zero Data
- cyberge99 1y agoWhat could go wrong?
- magicalhippo 1y agoJust don't hook it into the nuclear missile controls. We've seen[1] how that goes[2]. [1]: https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project [2]: https://en.wikipedia.org/wiki/The_Terminator https://en.wikipedia.org/wiki/The_Terminator
- koakuma-chan 1y ago[3] https://en.wikipedia.org/wiki/Re:Zero https://en.wikipedia.org/wiki/Re:Zero
- jasonjmcghee 1y agoConceptually, it's effectively a GAN
- magicalhippo 1y agoFor those not in the know, that's Generative Adversarial Networks[1], where two neural networks are trained in a competitive way. One network typically generates tasks for the other, and is rewarded if it manages to make the other network fail the task. The other network is rewarded if it successfully completes the task. Thus the adversarial network tries to find weaknesses to exploit, and the combined training makes the solving network much stronger. Or at least that's the idea. [1]: https://en.wikipedia.org/wiki/Generative_adversarial_network https://en.wikipedia.org/wiki/Generative_adversarial_network
- torginus 1y agoGAN's are a supervised training method, not really self-improving (after converging to being able to reproduce the training set).
- frumiousirc 1y agoMy initial thought as well. But, what is the "Discriminator" here? What grounds the training toward reality? The "Challenger" and "Solver" adversity alone can only serve to amplify hallucination. Ahh, GPT-4o is the arbiter. So, basically, this is a way to perform LLM model compression (GPT-4o to qwen3) while maximizing the in-distribution domain size. As such, it seems reasonable and useful. However the reliance on an arbiter LLM makes the claim that it will overcome the problem of a lack of training data unreasonable. Once the target LLM is scaled up to reach the in-distribution domain size of the arbiter, it seems to me it will turn back into a hallucination amplifier.
- djoldman 1y agoSee Figure 2. The solver/challenger is the GAN discriminator/generator. The challenger is trained to create difficult questions. The solver is trained to strengthen pathways that correctly solve the questions like so: > To guide the Challenger toward producing challenging yet solvable questions, we first define an uncertainty score. For a generated question x, we query the current Solver... The most frequent response is treated as the pseudo-label y˜(x), and we compute the Solver’s empirical accuracy....The uncertainty reward is then defined.... This function incentivizes questions where the Solver is maximally uncertain (accuracy approaches 50%) Identifying the best pseudo-label seems like it would be the limitation of the approach.
- frumiousirc 1y ago> Identifying the best pseudo-label seems like it would be the limitation of the approach. Yes, I think this says in a different way what I'm trying to express. In GAN, the Discriminator pegs the training to some chosen reality (assuming the "real" data set is truly real). In Challenger/Solver alone, there is no peg. The Solver could hallucinate consistently and "win" the race. It's the consistency that is the goal. With GPT-4o as an arbiter of the Challenger/Solver training it provides the reality peg (or rather, the peg that biases toward GPT-4o's training set).
- thom 1y agoFor values of zero quite far above zero.
- falcor84 1y agoWhat am I missing? From my skimming, there's zero external data beyond what is needed for the Challenger to generate questions.
- thom 1y agoAn existing trained LLM is an enormous amount of 'data' however it might be encoded. AlphaZero didn't start with Stockfish or a database of games.
- magicalhippo 1y agoAs I understand it the point of the article isn't to train a LLM from scratch, it's to teach a non-reasoning model to reason without additional explicit training data.
- YeGoblynQueenne 1y agoThe abstract does use the term "from scratch": >> To overcome this limitation, we introduce R-Zero, a fully autonomous framework that generates its own training data from scratch. Giving the benefit of the doubt, they're just using it wrong, but the way they use it sure reads like they claim they found a way to initialise LLMs with 0 data. Only the absurdity of the claim protects the reader from such misunderstanding, and that's never a good thing in a research paper.
- magicalhippo 1y agoIf you included the previous and following sentences, it's at least to me clear what they mean: However, existing methods for training such models still rely heavily on vast human-curated tasks and labels, typically via fine-tuning or reinforcement learning, which poses a fundamental bottleneck to advancing AI systems toward capabilities beyond human intelligence To overcome this limitation, we introduce R-Zero, a fully autonomous framework that generates its own training data from scratch. Starting from a single base LLM, R-Zero initializes two independent models with distinct roles, a Challenger and a Solver. Training a LLM is a multi-stage process[1], and they're tackling the stage at the end. That's where you do fine-tuning or reinforcement learning. They're not training a LLM from scratch. They're explicitly stating they start from a base LLM, ie a pretrained non-tuned model. As I understand it, and as they mention, training data for the latter stages has typically required high-quality human-curated samples in large numbers, even if they're augmented using LLMs, say by generating multiple variations of each human-curated training sample. Their proposal is to have a generative adversarial network generate that data without any initial human input, ie from scratch. [1]: https://snorkel.ai/blog/large-language-model-training-three-phases-shape-llm-training/ https://snorkel.ai/blog/large-language-model-training-three-...
- nakamoto_damacy 1y agoPerpetual Motion Machines were a thing at some point, too.
- YeGoblynQueenne 1y agoDon't laugh. PMMs work! I built mine ten years ago when I realised I could improve the SOTA by a huge 20%. I've been improving it for the last 10 years and I get an average performance boost of ~0.25 every year. We will have Free Energy in the next 10 years.
- ojo-rojo 1y agoI find your comment interesting, even though I'm not sure if I really get what you're saying. You built a perpetual motion machine? You then made improvements? Can you share details?
- suprfsat 1y agoGood news everyone, you've passed the Turing test.
- amelius 1y agoHmm, I guess I didn't pass it then.
- pas 1y agothey are claiming that they built a PMM prototype, which is not fully satisfying the business requirements yet, but they are on track to do so based on all the amazing documented validated peer-reviewed published progress they made already over the years!
- YeGoblynQueenne 1y agoThat!
- 1y ago
- clbrmbr 1y agoTerrible choice of name. DeepSeek developed a historically important model called “R-Zero” (this was the predecessor to R1 that was training without any coldstart SFT, and was very strong but difficult to read chain of thought because it code switches into Chinese and has no line breaks).
- neuroelectron 1y agoNow gamify it.
- Iv 1y ago"Starting from a single base LLM" Ok, zero data, except the data used in the teacher model.
- nickpsecurity 1y agoOnly 1-15TB of data processed at $10k-$100m depending on model size. Then, this shaves off a few hundred to a few grand on fine-tuning. I mean, we're still saving money at least.
- markmoscov 1y ago[dead]
- Davidzheng 1y agoI think in formal domain like lean it should actually be possible to do it from zero--but seems like no major successes no far
- freejazz 1y agoI still don't understand what a "reasoning" LLM is
- cluckindan 1y agoIt’s an LLM that has been trained and prompted to make users believe that the model is using logical reasoning to arrive at its output, when it is in fact still predicting the possible next output tokens, just like any other LLM. There may be additional feedback loops, but fundamentally, that is what it is doing. Sure, it will show you what steps it takes to arrive at a conclusion, but it is just predicting the steps, the conclusion and the potential validity of the aforementioned based on its training data, not actually evaluating the logic or the truthiness of the output. If you don’t believe me, ask your ”reasoning” LLM this question: What’s the name of the paternal great-great-grandfather of the son of Jacob’s son’s son’s son?
- sindriava 1y agoI won't read this because you're not really thinking, just pressing keyboard keys.
- cluckindan 1y agoJoke’s on you, I dictated it.
- sindriava 1y agoRich coming from the guy who moved his muscles until sounds came out. Also next time you should bother to at least copy paste your questions into any recent LLM, since they can all solve it without issue. But hallucations like this are common with non-reasoning HN users.
- cluckindan 1y agoBut can they solve it without referring to the Bible, or without mentioning anyone in the biblical Jacob’s family tree? Don’t think so. Humans solve that puzzle in a very different way than LLMs ”reason” about it.
- lawlessone 1y agoOK but how do you ensure it's improving in a direction that aligns with reality?