3 ms·
and get high token bandwidth?
by formvoltron 2mo ago
and get high token bandwidth?
- colordrops 2mo agoSimilar to a spark, which isn't blazing fast but usable.
- numpad0 2mo agoNot badly so because MoE models(identifiable by "CoolName-xxxB-AxxB" naming scheme) have bunch of branches in the middle that only one out of all gets non-zero values. Each of branches aka "Experts" as well as top/bottom parts are significantly smaller than the whole, and so CPU emulation of CUDA operations mixed with GPU taking as much as possible become not so out of question, unlike for dense models("CoolName-xxxB" without "-AxxB")