3 ms·
Per the screenshot, this is a DeepSeek running on a 192GB M2 Studio https://nitter.poast.org/ggerganov/status/1884612770093842728#m https://nitter.poast.org/gg
by un_ess 2y ago
Per the screenshot, this is a DeepSeek running on a 192GB M2 Studio
https://nitter.poast.org/ggerganov/status/1884612770093842728#m https://nitter.poast.org/ggerganov/status/188461277009384272...
The same on Nvidia (various models)
https://github.com/ggerganov/llama.cpp/issues/11474 https://github.com/ggerganov/llama.cpp/issues/11474
[1] this is a the model: https://huggingface.co/unsloth/DeepSeek-R1-GGUF/tree/main/DeepSeek-R1-UD-IQ1_S https://huggingface.co/unsloth/DeepSeek-R1-GGUF/tree/main/De...
- diggan 2y agoSo Apple M2 Studio does ~15 tks/second and A100-SXM4-80GB does 9 tks/second? I'm not sure I'm reading the results wrong or missing some vital context, but that sounds unlikely to me.
- achierius 2y agoThe studio has a lot more ram available to the GPU (up to 192gb) than the a100 (80gb), and iirc at least comparable memory bandwidth -- those are what matter when you're doing LLM inference, so the studio tends to win out there. Where the a100 and other similar chips dominate is in training &c, which is mostly a question of flops.
- diggan 2y ago> and iirc at least comparable memory bandwidth I don't think they do. From Wikipedia: > the M2 Pro, M2 Max, and M2 Ultra have approximately 200 GB/s, 400 GB/s, and 800 GB/s respectively From techpowerup: > NVIDIA A100 SXM4 80 GB - Memory bandwidth - 2.04 TB/s Seems to be a magnitude of difference, and that's just the bandwidth.