3 ms·Avatarl: Training language models from scratch with pure reinforcement learning2 points by haneefmubarak 1y ago