3 ms·One battle after another: using RL-guided reasoning for next-token prediction1 points by macleginn 1y ago