3 ms·
That's super helpful, thanks! I assume MoE is still prohibitively expensive to train, which is why we're not seeing massive MoE models?
by bytefactory 3y ago
That's super helpful, thanks!
I assume MoE is still prohibitively expensive to train, which is why we're not seeing massive MoE models?
- famouswaffles 3y agoNo. MoE models are far cheaper to train and far cheaper for inference. We're not seeing massive MoE models because they've typically well underperformed their dense counterparts. Only recently has it looked like we could get equitable performance from MoE architectures. https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2308.00951 https://arxiv.org/abs/2308.00951 In the first paper, you can see the underperformance i'm talking about. Flan-Moe-32b(259b total) scores 25.5% on MMLU pre Instruct tuning and 65.4 after. Flan 62b scores 55% before Instruct tuning and 59% after.
- bytefactory 3y agoFascinating. Thanks again!