2 ms·
> What they propose is effectively an MoE with separately learned routing as far as I can tell, with the heads of the LLM being task specific and there being a
by carbocation 3y ago
> What they propose is effectively an MoE with separately learned routing as far as I can tell, with the heads of the LLM being task specific and there being a task mapper to assign training labels.
What you are describing would indeed be an MoE model, but that's not where it ends. The manuscript continues [1]: "all agents become identical after all tasks have been learned and shared, and they all can master all tasks."
That is a substantial divergence from the MoE model!
To the extent that there is an MoE-like stage in their training, I think it's odd that the term "mixture of experts" is not mentioned nor is the literature on the topic cited.
1 = https://openreview.net/pdf?id=Jjl2c8kWUc https://openreview.net/pdf?id=Jjl2c8kWUc
- andreyk 3y agoGood point - it's a loose analogy training wise, and not applicable to inference time.