3 ms·
This is true for the pre-training step. What if advancements in the reinforcement learning steps performed later may benefit from more compute and more models p
by antirez 1y ago
This is true for the pre-training step. What if advancements in the reinforcement learning steps performed later may benefit from more compute and more models parameters? If right now the RL steps only help with sampling, that is, they only optimize the output of a given possible reply instead of the other (there are papers pointing at this: that if you generate many replies with just the common sampling methods, and you can verify correctness of the reply, then you discover that what RL helps with is selecting what was already potentially within the model output) this would be futile. But maybe advancements in the RL will do to LLMs what AlphaZero-alike models did with Chess/Go.
- cs702 1y agoIt's possible. We're talking about pretraining meaningfully larger models past the point at which they plateau, only to see if they can improve beyond that plateau with RL. Call it option (3). No one knows if it would work, and it would be very expensive, so only the largest players can try it, but why the heck not?