3 ms·
Apparently Kimi K3 has 104B parameters active at a time. So at 4 bits you'd need 52GB just to hold the active params. That said, in theory this same technique
by AussieWog93 2mo ago
Apparently Kimi K3 has 104B parameters active at a time. So at 4 bits you'd need 52GB just to hold the active params.
That said, in theory this same technique should be able to run it on a 64GB Macbook, probably at <1 tps.
- zozbot234 2mo agoOnly the sparse experts are 4-bit in native precision, and those take up ~25GB of active params. The dense parameters' native footprint is ~115GB. So in order to infer that natively on a 16GB machine you'd need to reload around ~135GB from disk at every token, which will take around 20.5 seconds at maximum 6.6 GB/s reading speed. This gives you a maximum theoretical performance of 176 tok/hr or 4224 tok/day when inferring at native precision. (Batching would be highly effective in aggregate since the bulk of what you're reloading is dense parameters, but your speed for any single session would still go down somewhat.) Of course all bets are off if you quantize the model highly; people are finding ways of fitting the whole thing in less than 600GB using extreme Q1 quants. Mind you, the outlook for a 64GB RAM machine isn't that different. You'd get a faster SSD (around 2.2x performance) and be able to cache more of your dense params. So your performance would probably be around 4x compared to the 16GB case.