3 ms·
The goal of this auxiliary module is to have the model recurse pre-emit trained on good thinking traces. This includes a 6-to-1 compression of thinking tokens.
by nmitchko 2mo ago
The goal of this auxiliary module is to have the model recurse pre-emit trained on good thinking traces. This includes a 6-to-1 compression of thinking tokens.
Therefore output tokens are decodeable, but are trained compressed. So they are approximations of faster thinking.
Interpretability is a mixed bag even with trained tools on top of existing models.