3 ms·
How does one effectively use something like this locally with consumer-grade hardware?
by aliljet 11mo ago
How does one effectively use something like this locally with consumer-grade hardware?
- tintor 11mo agoConsumer-grade hardware? Even at 4bits per param you would need 500GB of GPU VRAM just to load the weights. You also need VRAM for KV cache.
- CamperBob2 11mo agoIt's MoE-based, so you don't need that much VRAM. Nice if you can get it, of course.
- oceansweep 11mo agoEpyc Genoa CPU/Mobo + 700GB of DDR5 ram. The model is a MoE, so you don't need to stuff it all into VRAM, you can use a single 3090/5090 to hold the activated weights, and hold the remaining weights in DDR5 ram. Can see their deployment guide for reference here: https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/Kimi-K2-Thinking.md https://github.com/kvcache-ai/ktransformers/blob/main/doc/en...
- simonw 11mo agoOnce the MLX community get their teeth into it you might be able to run it on two 512GB M3 Ultra Mac Studios wired together - those are about $10,000 each though so that would be $20,000 total. Update: https://huggingface.co/mlx-community/Kimi-K2-Thinking https://huggingface.co/mlx-community/Kimi-K2-Thinking - and here it is running on two M3 Ultras: https://x.com/awnihannun/status/1986601104130646266 https://x.com/awnihannun/status/1986601104130646266