3 ms·
The MoE architecture allows you to keep the entire active model on a single GPU. If two consecutive tokens use the same export then the second token is going to
by hmottestad 2y ago
The MoE architecture allows you to keep the entire active model on a single GPU. If two consecutive tokens use the same export then the second token is going to be much faster.
- utopcell 2y agoWhat is the probability of that happening?
- zamadatix 2y agoDeepSeek V3/R1 uses 8 routed experts out of 256, so not all as often as one would like. That said, having even just a single GPU will greatly speed up prompt processing which is worth it even if the inference speed was the same. Ktransformers has a document about using CPU + a single 4090D to reach decent tokens/s but I'm not sure how much of the perf is due to the 4090D vs other optimizations/changes for the CPU side https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/DeepseekR1_V3_tutorial.md https://github.com/kvcache-ai/ktransformers/blob/main/doc/en... The final step of going to 6 experts instead of 8 feels like cheating (not a lossless optimization).
- genewitch 2y agowhere does 256 come from? it's repeated in here and elsewhere that a single expert is 37B sized, so you'd have to have way more than "several hundred billion parameters", to hold 256 of those? Maybe i don't understand the architecture, but if that's the case, then everyone repeating 37B doesn't, either.
- zamadatix 2y agoI think this diagram from the DeepSeekMoE paper explains it the clearest: https://i.imgur.com/CRKttob.png https://i.imgur.com/CRKttob.png The one on the right is how the feed forward layers of DeepSeek V3/R1 work, blue and green are experts, and everything in that right section is what counts as "active parameters". K (K=8 for these models, but you can customize that if you want) experts of 256 per layer are activated at a time. The 256 comes from the model file, it's just how many they chose to build it with. In these models there is also 1 shared expert which is always active in the layer. The router picks which k routed experts to use each forward pass and then a gating mechanism combines the outputs. If you sum the 1 shared expert + K routed experts + router + output networks you end up with 37 B parameters active for each feed forward layer pass. The individual experts are therefore much smaller than the total (probably something like 4 B parameters each? I've never really checked that directly). Or, for the short answer: "37 B is the active parameters of 9 experts + 'overhead', not the parameters of a single expert".
- genewitch 2y agoI understand all that, I am talking about a separate feature that is possibly backported or from llama.cpp. Where you have a small model that runs first and that is checked by a large model. I've seen 30%+ speedups using like 1.5B in front of a 15B for example. Two GPUs or more mean you can start to "keep" one or more of the experts hot on a GPU as well.