4 ms·
Could be the first open model to reach GPT-4 levels? Can't wait to see results of independant systematic human llm evaluation, it will surely take the first pla
by singularity2001 3y ago
Could be the first open model to reach GPT-4 levels? Can't wait to see results of independant systematic human llm evaluation, it will surely take the first place here:
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
Can it be compressed to run on mac studios?
- slowmovintarget 3y agoIt's very likely GPT-4 is an ensemble. A single model won't be able to keep up, even with this level of parameters. Run a fleet of these together, however...
- famouswaffles 3y agoIf the rumors are true, GPT-4 is a Sparse Mixture of Experts, not an ensemble.
- sunshadow 3y agoMixture of Experts is actually some sort of ensembling
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- bytefactory 3y agoCould you explain the difference to a noob, please?
- famouswaffles 3y agoAn ensemble is basically a mix of models for different tasks. One model could be an LLM, another image understanding model etc. Different tasks could be passed to different models or every task could be passed to all models and a result collated etc. MoE is...well basically a way to have a large model without computing all the parameters at once. So you take several smaller language models and you train them all on subsets of the same dataset. Then you train them to make predictions together. You could train for switching experts at the token level i.e one expert picks one token and another picks the next etc The "experts" are not clearly delineated or known. One "expert" could be a capital letter expert etc. People see GPT-4 being MoE and they go "Oh so questions about medicine are being passed to a separate model than questions about say Mathematics etc" but that's a misconception.
- bytefactory 3y agoThat's super helpful, thanks! I assume MoE is still prohibitively expensive to train, which is why we're not seeing massive MoE models?
- famouswaffles 3y agoNo. MoE models are far cheaper to train and far cheaper for inference. We're not seeing massive MoE models because they've typically well underperformed their dense counterparts. Only recently has it looked like we could get equitable performance from MoE architectures. https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2308.00951 https://arxiv.org/abs/2308.00951 In the first paper, you can see the underperformance i'm talking about. Flan-Moe-32b(259b total) scores 25.5% on MMLU pre Instruct tuning and 65.4 after. Flan 62b scores 55% before Instruct tuning and 59% after.
- bytefactory 3y agoFascinating. Thanks again!
- slowmovintarget 3y agoThank you for the correction.