3 ms·
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
by kamranjon 2mo ago
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
- nmitchko 2mo agoIt’s an adaptation of CoLaR, but my implementation is a little different: - Dedicated stop head to fire when latent thinking hits threshold - MTP support with training taking draft support as first class - different architectural layer 35 -> layer 42 writeback. So latents skip roughly 6.2 tokens of reasoning per token, then never make it to decoded output