3 ms·
I'm a bit confused as to why a Mixture of Experts (MOE) isn't one of the comparators. That seems like the most relevant direct comparator, rather than the sever
by carbocation 3y ago
I'm a bit confused as to why a Mixture of Experts (MOE) isn't one of the comparators. That seems like the most relevant direct comparator, rather than the several other paradigms that they cited.
- andreyk 3y agoMoE is a type of neural network architecture, not an approach to lifelong learning. As they say " We propose a new Shared Knowledge Lifelong Learning (SKILL) challenge, which deploys a decentralized population of LL agents that each sequentially learn different tasks, with all agents operating independently and in parallel." What they propose is effectively an MoE with separately learned routing as far as I can tell, with the heads of the LLM being task specific and there being a task mapper to assign training labels. The baselines based on Parameter-Isolation methods are sort of similar MoEs, in that additional weights are added and trained on each tasks (somewhat as experts would).
- 6gvONxR4sf7o 3y agoThere are network architectures for MoE, but afaik the concept of MoE is separate from them. MoE at least significantly predates the last decade's neural network boom.
- andreyk 3y agoafaik the term MoE is generally used in the context of neural networks, with the term mixture models being more general. But you are right it's not exclusive to NNs.
- carbocation 3y ago> What they propose is effectively an MoE with separately learned routing as far as I can tell, with the heads of the LLM being task specific and there being a task mapper to assign training labels. What you are describing would indeed be an MoE model, but that's not where it ends. The manuscript continues [1]: "all agents become identical after all tasks have been learned and shared, and they all can master all tasks." That is a substantial divergence from the MoE model! To the extent that there is an MoE-like stage in their training, I think it's odd that the term "mixture of experts" is not mentioned nor is the literature on the topic cited. 1 = https://openreview.net/pdf?id=Jjl2c8kWUc https://openreview.net/pdf?id=Jjl2c8kWUc
- andreyk 3y agoGood point - it's a loose analogy training wise, and not applicable to inference time.