2 ms·Digital Agent outperforms o1 by 15% – trained with new RL-variant similar to R111 points by let_tim_cook_ 2y ago