4 ms·
All parameters still need to be loaded into vram, it'll dynamically select two submodels to run on each token so it would be extremely slow to swap them out.
by andersa 3y ago
All parameters still need to be loaded into vram, it'll dynamically select two submodels to run on each token so it would be extremely slow to swap them out.
- behnamoh 3y agoThen what's the advantage of this technique compared to running a +50B model in the first place?
- nulld3v 3y agoThe model is quicker to evaluate. So quicker responses and more throughput.
- hnuser123456 3y agoSpeed, assuming you have the RAM to have it all loaded, it's faster than a fully connected network of the same size by 4x
- rileyphone 3y agoIt's better if you're hosting inference, worse if you are using it for a dedicated purpose. Presumably in the future it might make sense to share one local MoE among the different programs that use it, especially for a demand-heavy application like programming.
- andersa 3y agoIt's faster to run inference and training. Less memory bandwidth needed.
- irthomasthomas 3y agoI'm confused, though, how then do they claim it needs the same resources as 13B? Is that amortised over parallel usage or something?
- gorbypark 3y agoThe same compute resources, but not the same VRAM. It will more or less get you the same tokens per second as a ~13B model but should have significantly "higher quality output" than a single 13B model.