3 ms·
Could you explain the difference to a noob, please?
by bytefactory 3y ago
Could you explain the difference to a noob, please?
- famouswaffles 3y agoAn ensemble is basically a mix of models for different tasks. One model could be an LLM, another image understanding model etc. Different tasks could be passed to different models or every task could be passed to all models and a result collated etc. MoE is...well basically a way to have a large model without computing all the parameters at once. So you take several smaller language models and you train them all on subsets of the same dataset. Then you train them to make predictions together. You could train for switching experts at the token level i.e one expert picks one token and another picks the next etc The "experts" are not clearly delineated or known. One "expert" could be a capital letter expert etc. People see GPT-4 being MoE and they go "Oh so questions about medicine are being passed to a separate model than questions about say Mathematics etc" but that's a misconception.
- bytefactory 3y agoThat's super helpful, thanks! I assume MoE is still prohibitively expensive to train, which is why we're not seeing massive MoE models?
- famouswaffles 3y agoNo. MoE models are far cheaper to train and far cheaper for inference. We're not seeing massive MoE models because they've typically well underperformed their dense counterparts. Only recently has it looked like we could get equitable performance from MoE architectures. https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2308.00951 https://arxiv.org/abs/2308.00951 In the first paper, you can see the underperformance i'm talking about. Flan-Moe-32b(259b total) scores 25.5% on MMLU pre Instruct tuning and 65.4 after. Flan 62b scores 55% before Instruct tuning and 59% after.
- bytefactory 3y agoFascinating. Thanks again!