3 ms·
> if a llm will run with usable performance at that scale? Yes. The reason: MoE. They are able to run at a good speed because they don't load all of the weigh
by espadrine 2y ago
> if a llm will run with usable performance at that scale?
Yes.
The reason: MoE.
They are able to run at a good speed because they don't load all of the weights into the GPU cores.
For instance, DeepSeek R1 uses 404 GB in Q4 quantization[0], containing 256 experts of which 8 are routed to[1] (very roughly 13 GB per forward pass). With a memory bandwidth of 800 GB/s[3], the Mac Studio will be able to output 800/13 = 62 tokens per second.
[0]: https://ollama.com/library/deepseek-r1 https://ollama.com/library/deepseek-r1
[1]: https://arxiv.org/pdf/2412.19437 https://arxiv.org/pdf/2412.19437
[2]: https://www.apple.com/newsroom/2025/03/apple-unveils-new-mac-studio-the-most-powerful-mac-ever/ https://www.apple.com/newsroom/2025/03/apple-unveils-new-mac...
- fullstackchris 2y agoYou seem like you know what you are talking about... mind if I ask what your thoughts on quantization are? Its unclear to me if quantization affects quality... I feel like I've heard yes and no arguments
- sosuke 2y agoI don't know what I'm talking about but when I first asked your question this https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9 https://gist.github.com/Artefact2/b5f810600771265fc1e3944228... helped start me on a path to understanding. I think. But if you don't already know the question your asking is not at all something I could distill down into a sentence or to that would make sense to a lay-person. Even then I know I couldn't distill it at all sorry. Edit: I found this link I referenced above on quantized models by bartowski on huggingface https://huggingface.co/bartowski/Qwen2.5-Coder-14B-GGUF#which-file-should-i-choose https://huggingface.co/bartowski/Qwen2.5-Coder-14B-GGUF#whic...
- espadrine 2y agoThere is no question that quantization degrades quality. The GGUF R1 uses Q4_K_M, which, on Llama-3-8B, increases the perplexity by 0.18[0]. Many plots show increasing degradation as you quantize more[1]. That said, it is possible to train a model in a quantization-aware way[2][3], which improves the quality a bit, although not higher than the raw model. Also, a loss in quality may not be perceptible in a specific use-case. Famously LMArena.ai tested Llama 3.1 405B with bf16 and fp8, and the latter was only 2 Elo points below, well within measurement error. [0]: https://github.com/ggml-org/llama.cpp/blob/master/examples/quantize/quantize.cpp#L46 https://github.com/ggml-org/llama.cpp/blob/master/examples/q... [1]: https://github.com/ggml-org/llama.cpp/discussions/5063#discussioncomment-10906969 https://github.com/ggml-org/llama.cpp/discussions/5063#discu... [2]: https://pytorch.org/blog/quantization-aware-training/ https://pytorch.org/blog/quantization-aware-training/ [3]: https://mistral.ai/news/ministraux https://mistral.ai/news/ministraux
- Ambix 2y agoI did my own experiments and it looks like (surprisingly) Q4KM models often outperforms Q6 and Q8 quantised models. For bigger models (in range of 8B - 70B) the Q4KM is very good, there are no any degradation compared to full FP16 models.
- _aavaa_ 2y agoThis doesn’t sound correct. You don’t know which expert you’ll need for each layer, so you either keep them all loaded in memory or stream them from disk