3 ms·
LLaMA 3 with >=70B params will be launching this year, so I don't think this is something that will hold for long. And Mixtral 8x7B is a 56GB model, sparsely. F
by extheat 3y ago
LLaMA 3 with >=70B params will be launching this year, so I don't think this is something that will hold for long. And Mixtral 8x7B is a 56GB model, sparsely. For now I agree, for many companies it doesn't make sense to open source something you intend to sell for commercial use, so the biggest models will likely be withheld. However, the important more thing is that there is some open source model, whether it be from Meta or someone else, that can rival the best open source models. And it's not like the param count can literally go to infinity, there's going to be an upper bound that today's hardware can achieve.
- lhl 3y agoJust an FYI, Mixtral is a Sparse Mixture of Experts that has 47B parameters for memory costs (but 13B active parameters per token). For those interested in reading more about how it works: https://arxiv.org/pdf/2401.04088.pdf https://arxiv.org/pdf/2401.04088.pdf For those interested in some of the recent MoE work going on, some groups have been doing their own MoE adaptations, like this one, Sparsetral - this is pretty exciting as it's basically an MoE LoRA implementation that runs a 16x7B at 9.4B total parameters (the original paper introduced a model, Camelidae-8x34B, that ran at 38B total parameters, 35B activated parameters). For those interested, best to start here for discussion and links: https://www.reddit.com/r/LocalLLaMA/comments/1ajwijf/model_release_sparsetral/ https://www.reddit.com/r/LocalLLaMA/comments/1ajwijf/model_r...