4 ms·
Their test time RL approach seems a bit fishy. From what I understand, TTRL works by asking a language model to generate simpler versions of the test case. Once
by barteloniu 2y ago
Their test time RL approach seems a bit fishy. From what I understand, TTRL works by asking a language model to generate simpler versions of the test case. Once we have the simpler problems, we run RL on them, hoping that an improvement on the simplified cases will also strengthen the model performance on the original problem.
The issue is, they use a numerical integrator to verify the simpler problems. One could imagine a scenario where a barely simpler problem is generated, and the model is allowed to train on pretty much the test case knowing the ground truth. Seems like training on the test set.
The rest of the paper is nice though.
- thomasahle 2y ago> the model is allowed to train on pretty much the test case knowing the ground truth The task is to solve the integral symbolically, though, right? It's a hard problem to solve, even if the model is given access to a numerical integrator tool it can use on the main problem itself.
- barteloniu 2y agoThat's a fair point.