4 ms·
Thats not really correct. It's actually the beginning of test time scaling. R1 has shown that a very simple reinforcement learning scheme can be used to teach
by cpldcpu 2y ago
Thats not really correct.
It's actually the beginning of test time scaling. R1 has shown that a very simple reinforcement learning scheme can be used to teach the model how to think in a chain-of-though as an emergent property.
No addition pretraining data needed! Only more compute.