6 ms·
Cool result, but worth highlighting two points: - Model is finetuned from Qwen-2.5 Instruct, which includes millions of specially filtered math examples in bot
by highfrequency 2y ago
Cool result, but worth highlighting two points:
- Model is finetuned from Qwen-2.5 Instruct, which includes millions of specially filtered math examples in both pretraining and supervised fine-tuning already.
- To generate the perfect 817 math examples for LIMO, they used state of the art models like R1 to filter down from an initial pool of 10 million math problems. In other words, a whole lot of intelligence was used to craft a maximally informative and distilled set of fine-tuning data. It’s not very clear to me if this is more or less impressive than getting the same result by simply fine-tuning on the 10 million initial pool, but I suppose that would make for a worse headline.
- smallerize 2y agoYeah, but it's cheaper. The context right now is that OpenAI, with first-mover advantage, cutting-edge-hardware, and tens of billions of dollars of investment, are not getting benchmark performance better than Chinese-developed models that are trained with cut-down nvidia GPUs and a lot less money.
- rfoo 2y agoBut... they are? o3-mini is faster than DeepSeek-R1 and has comparable capability. And while I hate "AGI achieved internally" meme, o3 is significantly better than o1. Though I doubt how long until DeepSeek-R3 happens. They could skip R2 too citing Cloudflare R2 :P
- pama 2y agoA big part of why R1 is much slowerr than o3-mini is that inference optimization is not yet performed on most solutions for serving R1 models (so R1 is rather comparable to o1 or o1 pro in terms of latency rather than o1-mini or o3-mini). The MoE is already relatively efficient if perfectly load balanced in an inference setting and should have latencies and throughputs that are equal to or faster than equivalent dense models with 37B parameters. In practice due to MLA inference should be much faster yet for long contexts compared to typical dense models. If DeepSeek or someone else tried to distill the model onto another MoE architecture with even less active parameters and properly implement speculative decoding on top, one could gain additional speedups in inference. I imagine we will see these things but it takes a bit of time till they are all public.
- rfoo 2y agoI know that, I'm in this game. I was comparing API throughput/ttft/ttbt of DeekSeek's own R1 API before it went viral in the West, and o3-mini. I remain unconvinced that DeepSeek themselves didn't optimize their own V3 inference good enough and left another 2x~3x improvement on the table.
- pama 2y agoI am sure DeepSeek did optimize the inference cost of R1. They did not yet release an efficient MoE downscaling of it, ie an R1-mini.
- smallerize 2y agoI actually forgot that o3-mini was available now. I was using o1 numbers.
- rvnx 2y agoI think you could reconsider DeepSeek-R1: it's actually really good. In comparison, o3-mini gets very vague in its reasoning, and gives surprisingly unhelpful answers (getting too short). Plus, let's not forget, R1 is available to use and modify under MIT license, which is great.
- amingilani 2y agoWhy is everyone is so critical of using information from a previous model to make a more efficient model. There’s nothing wrong with making progress using prior work. And increasing efficiency is progress. You wouldn’t criticize someone’s kombucha because they didn’t piece their SCOBY (symbiotic culture of bacteria and yeast) together microbe by microbe.
- deleted 2y ago[deleted]
- carschno 2y agoYou are looking at it from a product perspective. From a scientific perspective, it just means the respective benchmark is meaningless, so we don't know how well such a model generalizes.
- EGreg 2y agoNot so! From a scientific perspective the result you can achieve matters, no one is a blank slate. For humans this is true as well. The way you teach matters. Look at how the bell curve got absolutely demolished for example when math was taught this way: https://archive.nytimes.com/opinionator.blogs.nytimes.com/2011/04/21/teaching-math-advanced-discussion/ https://archive.nytimes.com/opinionator.blogs.nytimes.com/20...
- h0l0cube 2y agoAnother way to look at this is: The first assembly language compiler was handcoded in binary to begin with, and then that compiler's machine code was translated to the more expressive language (assembly). Similar for Fortran/C/etc. from assembly code. Progressively, more expressive languages have been bootstrapped from prior lower-level languages. In a similar way, perhaps a more concise LLM can be built by utilizing a less efficient one?
- btown 2y agoThere is a valid criticism that when you rely heavily on synthetic outputs, you bring along the precursor model's biases and assumptions without fully knowing the limitations of the data set the precursor model was trained on, as well as intentional adjustments made by the designers of the precursor model to favor certain geopolitical goals. But that's not the criticism that I'm often seeing; it's more that there's an "unfair" amount of press coverage towards new models that rely, in the critics' views, more on distillation than on "true" innovation. It's worth noting that there are many parties with significant motivation to build public sympathy that only "true" innovation should be valued, and it is only their highly-valued investments that can uniquely execute in that space. Cutting-edge models built in caves with a box of their scraps are counter to that narrative. It's worth considering https://paulgraham.com/submarine.html https://paulgraham.com/submarine.html in this context, and understanding whether it is truly "everyone" that is critical in this way.
- armcat 2y agoYes, the authors explicitly highlighted those two points in the abstract, in terms of them being the elicitation threshold for complex reasoning, namely, an extremely complete pre-trained foundation model, and a set of extremely high quality examples post-training. To your question on finetuning on the initial 10 million pool - intuitively, it would require tremendous amount of finetuning data to move the needle - you really won't be able to move the gradients much with just 817 examples, that initial pool is effectively enforcing pretty rigid regularization. There is now an increasing interest in showing that small data with inference time scaling is providing significant yield. Couple of recent examples: * TinyZero: https://github.com/Jiayi-Pan/TinyZero https://github.com/Jiayi-Pan/TinyZero * s1 Simple Test Time Scaling: https://arxiv.org/abs/2501.19393 https://arxiv.org/abs/2501.19393
- highfrequency 2y agoThe abstract doesn’t specify that the 857 training examples were filtered down by R1 from 10 million initial questions. This helps to understand the result better: it is in large part a testament to R1 and similar models’ remarkable ability sift through and identify/construct perfect training data for other models.
- Eisenstein 2y agoIsn't every progression in technology a result of the previous advance in technology enabling it?
- highfrequency 2y agoYes, but these three types of progress are worth distinguishing: 1. Mt. Everest is summited for the first time. 2. An easier or more direct route to the Everest summit is discovered. 3. Someone finds that if a more experienced climber is already at the summit and drops down a series of rope ladders and oxygen tanks and cabins at key points, then it is even easier to make the summit because you can now pack lighter. All three are interesting, worth discussing etc. But it would be a bit of a stretch to conclude from the third one that “less is more” because you don’t need to bring so much gear when someone else brings it for you. For example, Attention is All You Need had a similar title. But the whole point was that they did not use recurrent networks at any stage in the learning process. My point is not to discredit this result but to frame it properly: reasoning models like R1/O1 are incredibly efficient at distilling knowledge to smaller non-reasoning models.
- trott 2y agoAnother way to look at this is that there are 12,290 bits of information in choosing 817 samples from 10,000,000.
- TOMDM 2y agoAnd much more information when selecting just as many examples from quadrillions of randomly generated examples. The information from the selection criteria isn't available to the model, just the chosen samples.
- orbital-decay 2y ago>In other words, a whole lot of intelligence was used to craft a maximally informative and distilled set of fine-tuning data. Sounds like any textbook. (and generally the process of knowledge compression over generations that made us who we are)
- yishanchuan 2y agoSure,but just in mathematical reasoning. If future works contain mathematical logic reasoning, it will perfect.
- EternalFury 2y agoJust imagine a textbook that gives you the understanding you need to score high in math competitions…and it describes less than 1,000 problems. This in itself is a major discovery in metacognition.
- robotresearcher 2y agoIt's one more textbook, not one textbook. I'm not knocking the work. They report large improvements using relatively little data. That's good. But let's be clear that this is further training of a good sized LLM that has read far, far more than any human that ever lived already.
- EternalFury 2y agoI know. The question is: How much of the Internet trove, including the smart bits, but also the tremendous amount of inane content, is actually useful to building the foundation that allows 1,000 problems to have such an effect?
- sdenton4 2y agoWell, there's this, which comes close: https://www.wiley.com/en-us/The+Art+and+Craft+of+Problem+Solving%2C+3rd+Edition-p-9781119239901 https://www.wiley.com/en-us/The+Art+and+Craft+of+Problem+Sol... Most of the math competitions people are working on are high school math competitions - these have problems from a relatively small set of mathematics, so that high school students can reasonably know the appropriate background.
- mattigames 2y agoYou are missing three point, it's about stating the importance of the preselection, now we know that we may not need huge amounts of data for similar results in other reasoning areas, only highly curated data, yes, sometimes by models themselves but not necessarily.
- Terretta 2y ago> To generate the perfect 817 math examples for LIMO, they used state of the art models like R1 to filter down from an initial pool of 10 million math problems. In other words, a whole lot of intelligence was used to craft a maximally informative and distilled set of fine-tuning data The paper, and this comment, seem awfully reminiscent of creating a textbook of curated "maximally informative and distilled" set of cognitive examples to teach students with foundational learning a next level of reasoning. The last few years of LLM progress have shown we can predict human "reasoning" responses to inputs by modeling likely human responses as if LLM generated. Put another way, most responses are not particularly reasoned, but chain of tokgen*. Sit near someone who "talks to herself" while doing problems and it's even more evident. --- * tokgen definition: Listen to conversations in a cafeteria. Many are something other than thoughtful, responses that follow the prompts, with near perfect predictability. To differentiate from these responses and speech that comes after a pause and reflect, one can use the labels thought versus token generation or tokgen.
- ciphix 2y agoAfter reviewing the paper and GitHub training dataset, I have the following observations: The 800+ training samples, each containing solutions with detailed reasoning steps, were primarily generated by DeepSeek r1 and advanced models. The reasoning processes within these training solutions are crucial. It's possible that the advanced models have encoded these reasoning processes through the generated samples. Given a sufficiently large model, it can effectively restore such reasoning weights, effectively adding a delta from DeepSeek r1, among others. Therefore, it's not surprising that, with relatively few fine-tuning data, Qwen 2.5 has achieved such significant improvements. This is merely a conjecture. Further research is needed to analyze and visualize the changes in network weights before and after fine-tuning.
- GTP 2y ago>The last few years of LLM progress have shown we can predict human "reasoning" responses to inputs by modeling likely human responses as if LLM generated. Put another way, most responses are not particularly reasoned, but chain of tokgen*. Sorry, but I don't get the point of your comment as a whole, and of this part in particular. Yes, most human day-to-day conversations are quite predictable, but some people are still capable of generating original thoughts from time to time. And still, how is it related to the comment you are replying to?