2 ms·
however, if you need to swap experts on each token, you might as well run on cpu.
by read_if_gay_ 3y ago
however, if you need to swap experts on each token, you might as well run on cpu.
- tarruda 3y ago> Presumably, the same expert would frequently be selected for a number of tokens in a row In other words, assuming you ask a coding question and there's a coding expert in the mix, it would answer it completely.
- read_if_gay_ 3y agoyes I read that. do you think it's reasonable to assume that the same expert will be selected so consistently that model swapping times won't dominate total runtime?
- tarruda 3y agoNo idea TBH, we'll have to wait and see. Some say it might be possible to efficiently swap the expert weights if you can fit everything in RAM: https://x.com/brandnarb/status/1733163321036075368?s=20 https://x.com/brandnarb/status/1733163321036075368?s=20
- ttul 3y agoSee my poorly educated answer above. I don’t think that’s how MoE actually works. A new mixture of experts is chosen for every new context.