2 ms·
Our primary focus is on RL post-training. We think that is the best way to get the model to be a strong interactive agent.
by srush 11mo ago
Our primary focus is on RL post-training. We think that is the best way to get the model to be a strong interactive agent.