3 ms·
That's a dense model. Of course it will do worse. Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it
by nacs 12d ago
That's a dense model. Of course it will do worse.
Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
- peri-cl 12d agoSurprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth, https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38flashnext_on_1x_rtx_5090_tg50_ts_pp2300_ts/ https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f... (Note it's a sparse MoE with only 6B active).
- nacs 12d agoGood to know thanks. That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
- well_ackshually 12d agoUnlike a 256GB M5 Ultra that is $10k+.
- nacs 12d agoApple product won't be the cheapest but it is a full package (CPU, RAM, VRAM/GPU, fast-storage, etc). If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).