3 ms·
M3 Ultra has a big GPU with 819 GB/sec bandwidth. LLM performance is twice as fast as RTX 5090 https://creativestrategies.com/mac-studio-m3-ultra-ai-workstati
by voidspark 1y ago
M3 Ultra has a big GPU with 819 GB/sec bandwidth.
LLM performance is twice as fast as RTX 5090
https://creativestrategies.com/mac-studio-m3-ultra-ai-workstation-review/ https://creativestrategies.com/mac-studio-m3-ultra-ai-workst...
- behnamoh 1y ago> LLM performance is twice as fast as RTX 5090 your tests are wrong. you used MLX for Mac Studio (optimized for Apple Silicon) but you didn't use vLLM for 5090. There's no way a machine with half the bandwidth of 5090 delivers twice as fast tok/s.
- seanmcdirmid 1y agoUnless it’s a large model that doesn’t fit in the 5090, bust that’s no longer a $4k macstudio I think.
- behnamoh 1y agothat's orthogonal to the speed discussion. also, the GP was mostly testing models that fit in both 5090 and Mac Studio.
- deleted 1y ago[deleted]
- voidspark 1y ago$4k will get you a 96 GB Mac Studio with M3 Ultra (819 GB/sec). That's 3x the RAM of the 5090.
- mdp2021 1y ago> That's 3x the RAM of the 5090 And a bit less than half the bandwidth (saying for completeness).
- voidspark 1y agoYeah that's probably wrong. But the M3 Ultra is good enough for local inferencing, in any case.