3 ms·
I think more likely it's for finetuning a pre-trained model like GPT-4, kinda like RLHF, but in this case using reinforcement learning somewhat similar to Alpha
by macrolime 3y ago
I think more likely it's for finetuning a pre-trained model like GPT-4, kinda like RLHF, but in this case using reinforcement learning somewhat similar to AlphaZero. The model gets pre-trained and then fine-tuned to achieve mastery in tasks like mathematics and programming, using something like what you say and probably something like tree of thought and some self reflection to generate the data that it's using reinforcement learning to improve on.
What you get then is a way to get a pre-trained model to keep practicing certain tasks like chess, go, math, programming and many other things as it gets figured out how to do it.
- wegfawefgawefg 3y agoI do not think that is correct as the RL in RLHF already stands for reinforcement learning. :^) However, I do think you are right that self play, and something like reinforcement learning will be involved more in the future of ML. Traditional "data-first" ml has limits. Tesla conceded to RL for parking lots, where the action and state space was too unknowable for hand designed heuristics to work well. In Deep Reinforcement Learning making a model just copy data is called "behavior cloning", and in every paper I have seen it results in considerably worse peak performance than letting the agent learn from its own efforts. Given that wisdom alone, we are under the performance ceiling with pure language models.