4 ms·
> One thing you could do with MoE is giving each expert different subsets of the input tokens. Don't MoE's route tokens to experts after the attention step? Th
by declaredapple 3y ago
> One thing you could do with MoE is giving each expert different subsets of the input tokens.
Don't MoE's route tokens to experts after the attention step? That wouldn't solve the n^2 issue the attention step has.
If you split the tokens before the attention step, that would mean those tokens would have no relationship to each other - it would be like inferring two prompts in parallel. That would defeat the point of a 10M context