3 ms·Scaling Reinforcement Learning for Trillion-Scale Thinking Model3 points by mountainview 11mo ago