3 ms·
First graph tells the story - below a certain model size (500m params), reinforcement learning is close to useless. Above this (task-dependent) model size thres
by highfrequency 2y ago
First graph tells the story - below a certain model size (500m params), reinforcement learning is close to useless. Above this (task-dependent) model size threshold, reinforcement learning basically works.
I suspect this is what we saw play out with math/coding reasoning models - until recently, the base models were not good enough for ~random output search to hit on a correct path with any reasonable frequency. Below this threshold of base model intelligence, the only efficient way forward was to collect plain supervised data (either through human labeled math problem solutions [1] or meticulous filtering of web text [2].
But as soon the base model (in this case Deepseek V3) breaks through and can actually solve a decent fraction of math problems, then reinforcement learning (plus other simple tricks like chain-of-thought prompting, simple ensemble voting, etc.) can easily juice the results through the following loop:
1) random search through different solution paths
2) identify the correct solution paths based on the final answer
3) train on the correct solution paths
The exciting thing is that not only can RL bump up the performance of the current base model, but it can be used to generate new high-quality reasoning trace data, which was in painfully short-supply for training the initial models. This leads to a new wave of base models with better one-pass intuition, which leads to more efficient reinforcement learning search on harder problems, which leads to better training data...
Note that this was basically impossible for non-LLM models in the past. You could always juice ImageNet classification performance with a simple ensemble of identically trained models, but that path didn't lead anywhere interesting because a juiced model didn't allow the creation of new synthetic data that was superior to the data it was trained on. The key difference is that LLMs not only output the solution but also output a solution path with all the intermediate steps - and these searched-and-filtered solution paths are much more valuable than the vast majority of the model's initial training data.
[1] https://arxiv.org/abs/2305.20050 https://arxiv.org/abs/2305.20050
[2] https://arxiv.org/abs/2402.03300 https://arxiv.org/abs/2402.03300 and https://arxiv.org/abs/2206.14858 https://arxiv.org/abs/2206.14858
- krackers 2y agoYeah that is the hypothesis some people have floated for as to why we only "just now" discovered this simple idea: https://news.ycombinator.com/item?id=42837349 https://news.ycombinator.com/item?id=42837349