4 ms·
For R1-Zero they did RL on two properties: a) there is a 'thinking' box and an 'answer box' (format constraint) b) the answer box has a correct answer Note t
by tmnvdb 2y ago
For R1-Zero they did RL on two properties:
a) there is a 'thinking' box and an 'answer box' (format constraint)
b) the answer box has a correct answer
Note that there is nothing about the contents of the thinking box in the reinforcement learning. Only that there is such a box.
They then observe that during RL, the model will start generating more and more stuff in the thinking box.
In essence this is the emergence of using more test-time compute to improve answers.
When reading out the thinking box, they found the model reflecting on its own answers, going back and changing its mind, and reflecting on the question, etc, similar to what we can see with R1 now. These are fully emergent phenomenon, without any prompting to do such a thing!
(The reasoning output would sometimes be a bit garbled and switch languages randomly, so for R1 they added some constraints to make the thinking content intelligble to humans. This actually made the model slightly worse at answering questions correctly.)
- mercer 2y agoWow. That's even cooler (and somehow simpler) than I thought it was.