3 ms·
Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottl
by peri-cl 5d ago
Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,
https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38flashnext_on_1x_rtx_5090_tg50_ts_pp2300_ts/ https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...
(Note it's a sparse MoE with only 6B active).
- nacs 5d agoGood to know thanks. That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
- well_ackshually 5d agoUnlike a 256GB M5 Ultra that is $10k+.
- nacs 5d agoApple product won't be the cheapest but it is a full package (CPU, RAM, VRAM/GPU, fast-storage, etc). If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).