2 ms·
We tried this, and it works :) https://arxiv.org/abs/2508.06111 https://arxiv.org/abs/2508.06111 You have to be careful about degenerate / duplicate Qs, as a s
by crs_gentleman 3mo ago
We tried this, and it works :) https://arxiv.org/abs/2508.06111 https://arxiv.org/abs/2508.06111
You have to be careful about degenerate / duplicate Qs, as a sibling commenter mentioned.
Recently though, we found that reasoning models have trouble making code-output-prediction tasks (the initial family of verifiable tasks we started with) which other reasoning models can't solve.
We started looking into harder / more agentic tasks (e.g. passing tests, using AISI's Inspect framework) but deprioritised.