3 ms·
Diverging from the sampler used in RL training is not good for long multi-turn results-- it's a great way to knock models into reasoning loops that wouldn't oth
by nullc 2mo ago
Diverging from the sampler used in RL training is not good for long multi-turn results-- it's a great way to knock models into reasoning loops that wouldn't otherwise.
- Der_Einzige 2mo agoPeer reviews NeurIPS caliber paper to prove that? Because I can show you one titled "Long context generation is a sampling problem"...
- WithinReason 2mo agoShow us, we're curious. Did you upload to ArXiv yet?