8 ms·
From the R1 paper In this study, we demonstrate that reasoning capabilities can be significantly improved through large-scale reinforcement learning (RL), even
by deepGem 2y ago
From the R1 paper
In this study, we demonstrate that reasoning capabilities can be significantly
improved through large-scale reinforcement learning (RL), even without using supervised
fine-tuning (SFT) as a cold start. Furthermore, performance can be further enhanced with
the inclusion of a small amount of cold-start data
Is this cold start data what OpenAI is claiming their output ? If so what's the big deal ?
- Imnimo 2y agoDeepSeek claims that the cold-start data is from DeepSeekV3, which is the model that has the $5.5M pricetag. If that data were actually the output of o1 (a model that had a much higher training cost, and its own RL post-training), that would significantly change the narrative of R1's development, and what's possible to build from scratch on a comparable training budget.
- TheGeminon 2y agoIn the paper DeepSeek just says they have ~800k responses that they used for the cold start data on R1, and are very vague about how they got it: > To collect such data, we have explored several approaches: using few-shot prompting with a long CoT as an example, directly prompting models to generate detailed answers with reflection and verification, gathering DeepSeek-R1-Zero outputs in a readable format, and refining the results through post-processing by human annotators.
- Imnimo 2y agoMy surface-level reading of these two sections is that the 800k samples come from R1-Zero (i.e. "the above RL training") and V3: >We curate reasoning prompts and generate reasoning trajectories by performing rejection sampling from the checkpoint from the above RL training. In the previous stage, we only included data that could be evaluated using rule-based rewards. However, in this stage, we expand the dataset by incorporating additional data, some of which use a generative reward model by feeding the ground-truth and model predictions into DeepSeek-V3 for judgment. >For non-reasoning data, such as writing, factual QA, self-cognition, and translation, we adopt the DeepSeek-V3 pipeline and reuse portions of the SFT dataset of DeepSeek-V3. For certain non-reasoning tasks, we call DeepSeek-V3 to generate a potential chain-of-thought before answering the question by prompting. The non-reasoning portion of the DeepSeek-V3 dataset is described as: >For non-reasoning data, such as creative writing, role-play, and simple question answering, we utilize DeepSeek-V2.5 to generate responses and enlist human annotators to verify the accuracy and correctness of the data. I think if we were to take them at their word on all this, it would imply there is no specific OpenAI data in their pipeline (other than perhaps their pretraining corpus containing some incidental ChatGPT outputs that are posted on the web). I guess it's unclear where they got the "reasoning prompts" and corresponding answers, so you could sneak in some OpenAI data there?
- deepGem 2y agoThat's what I am gathering as well. Where is OpenAI going to have substantial proof to claim that their outputs were used ? The reasoning prompts and answers for SFT from V3 you mean ? No idea. For that matter you have no idea where OpenAI got this data from either. If they open this can of worms, their can of worms will be opened as well.
- IAmGraydon 2y ago>Where is OpenAI going to have substantial proof to claim that their outputs were used ? I assume in their API logs.
- rekttrader 2y agoShibboleths in output data
- joe_the_user 2y agoIt's like the claim "they showed anyone create a powerful from scratch" becomes "false yet true". Maybe they needed OpenAI for their process. But now that their model is open source, anyone can use that as their cold start and spend the same amount. "From scratch" is a moving target. No one who makes their model with massive data from the net is really doing anything from scratch.
- bmicraft 2y agoYeah, but that kills the implied hope of building a better model for cheaper. Like this you'll always have a ceiling of being a bit worse then the openai models.
- reassess_blind 2y agoIsn't DeepSeek a bit better, not worse?
- roenxi 2y agoThe logic doesn't exactly hold, it is like saying that a student is limited by their teachers. It is certainly possible that a bad teacher will hold the student back, but ultimately a student can lag or improve on the teacher without only a little extra stimulus. They probably would need some other source of truth than an existing model, but it isn't clear how much additional data is needed.
- diedyesterday 2y agoDon't forget that this model probably has far less params than o1 or even 4o. This is a compression/distillation, which means it frees up so much compute resources to build models much powerful than o1. At least this allows further scaling compute-wise (if not in the amount of, non-synthetic, source material available for training).
- Loic 2y agoNot for me. As I build a chemical factory, I do not reinvent everything. They are using the current SOTA tools and models to build new models for cheaper.
- vlovich123 2y agoIf R1 were better than O1, yes you would be right. But the reporting I’ve seen is that it’s almost as good. Being able to copy cutting edge models won’t advance the state of the art in terms of intelligence. They have made improvements in other area, but if they reused O1 to train their model, that would be effectively a ctrl-c / ctrl-v strictly in terms of task performance.
- PeterStuer 2y agoStrong disagree. Copy/paste would mean they took o1's weights and started finetuning from there. That is ot what happened here at all.
- skinner_ 2y agoWhen you build a new model, there is a spectrum of how you use the old model: 1. taking the weights, 2. training on the logits, 3. training on model output, 4. training from scratch. We don't know how much advantage #3 gives. It might be the case that with enough output from the old model, it is almost as useful as taking the weights.
- vlovich123 2y agoFirst, there could have been industrial espionage involved so who knows. Ignoring that, you’re missing what I’m saying. Think of it this way - if it requires O1’s input to reach almost the same task performance, then this approach gives you a cheap way to replicate the performance of a leading edge model at a fraction of the cost. It does not give you a way to train something that beats a cutting edge model. Cutting edge models require a lot of R&D & capital expenditure - if they’re just going to be trivially copied after public availability, the response is going to be legislation to keep the incentive there to keep meaningful investments in that area. Otherwise you’re going to have another AI winter where progress shrivels because investment dollars dry up. That’s why it’s so hard to understand the true cost of training Deepseek whereas it’s a little bit easier for cutting edge models (& even then still difficult).
- powerapple 2y agoI lean on the idea that R1-Zero was trained from cold start, at the same time, they have tried many things including using OpenAI APIs. These things can happen in parallel.