3 ms·
Would this be a reasonable explanation? > MLPs are universal function approximators, but these models are big enough that it is better to train many small func
by DougBTX 2y ago
Would this be a reasonable explanation?
> MLPs are universal function approximators, but these models are big enough that it is better to train many small functions rather than a single unified function. MoE is a mechanism to force different parts of the model to learn distinct functions.
- samus 2y agoIt misses the crucial detail that every transformer layer chooses the experts independently from the others. Of course they still indirectly influence each other since each layer processes the output of the previous one.