2 ms·
It is MoE. It needs to engage multiple experts when the problem is complex or unclear. So you naturally see more of those simply as a primitive it learns to use
by mordae 2mo ago
It is MoE. It needs to engage multiple experts when the problem is complex or unclear. So you naturally see more of those simply as a primitive it learns to use to page in more diverse set of weights. Remember that each token is just 6 experts out of 256. So it literally needs to tell its router that it needs a different set the next time.
And this memory control primitive leaks into the reasoning chain, because it has no other channel for it available and we do not know how to train any other channel.
On the flip side, it tends to converge quickly, roughly proportional to the actual difficulty / clarity of the task.
- aesthesia 2mo agoHas this hypothesis that reasoning tokens help the model engage a wider range of experts been tested? I'm a little skeptical that it's the main driver of extended reasoning traces, particularly because MoE models are generally already trained so that expert activations are as uniformly distributed as possible. But it would be interesting to know to what extent this happens.