6 ms·
The mainstream options seem to be Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999 Nvidia DGX Spark, ~1000 tops fp4, 128GB RAM, $3999 Mac Studio max spe
by cherioo 1y ago
The mainstream options seem to be
Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999
Nvidia DGX Spark, ~1000 tops fp4, 128GB RAM, $3999
Mac Studio max spec, ~120 tflops (fp16?), 512GB RAM, 3x bandwidth, $9499
DGX Spark appears to potentially offer the most token per second, but less useful/value as everyday pc.
- aurareturn 1y agoMac Studio max spec, ~120 tflops (fp16?), 384GB RAM, 3x bandwidth, $9499 512GB. DGX has 256GB/s bandwidth so it wouldn't offer the most tokens/s.
- rz2k 1y agoPerhaps they are referring to default GPU allocation that is 75% of the unified memory, but it is trivial to increase it.
- jauntywundrkind 1y agoThe GPU memory allocation refers to how capacity is alloted, not bandwidth. Sounds like the same 256-bit/quad-channel 8000MHz lpddr5 you can get today with Strix Halo.
- rz2k 1y ago384GB is 75% of 512GB. The M3 Ultra bandwidth is over 800GB/s, though potentially less in practice. Using an M3 Ultra I think the performance is pretty remarkable for inference and concerns about prompt processing being slow in particular are greatly exaggerated. Maybe the advantage of the DGX Spark will be for training or fine tuning.
- vid 1y agoI very consistently see people say prompt processing is slow for larger context sizes ("notoriously slow"), something that is much less of an issue with eg CUDA setups.
- Art9681 1y agoDepends on the model. gpt-oss-120b will easily crunch large prompts in a few seconds. It's remarkable. It's gpt-4-mini at home.
- echelon 1y agotokens/s/$ then.
- deleted 1y ago[deleted]
- jauntywundrkind 1y agoNVidia Spark is $4000. Or, will be, supposedly whenever it comes out. Also notably, Strix Halo and DGX Spark are both ~275GBps memory bandwidth. Not always but in many machine learning cases it feels like that's going to be the limiting factor.
- UncleOxidant 1y ago> Ryzen AI Max 395+, ~120 tops (fp8?), 128GB RAM, $1999 Just got my Framework PC last week. It's easy to setup to run LLMs locally - you have to use Fedora 42, though, because it has the latest drivers. It was super easy to get qwen3-coder-30b (8 bit quant) running in LMStudio at 36 tok/sec.
- hasperdi 1y agoHi could you share if you get a decent coding performance (quality wise) with this setup? IE. Is it good enough to replace say Claude Code?
- UncleOxidant 1y agoqwen3-coder-30b is surprisingly good for a smallish model, but it's not going to replace Claude Code. Maybe if you're using it for Python it could do well enough. I've been trying it with C code generation and it's not bad, but certainly not at Claude Code level. I hope they come out with a qwen coder model in the 60b to 80b range - something like that would give higher quality results and likely still run in the 15 tok/sec range which would be usable.
- pixelpoet 1y agoVery encouraging result, I'm waiting super anxiously for mine! How much memory did you allocate for the iGPU?
- UncleOxidant 1y agoI haven't done any fiddling with that yet. Out of the box it seems to allocate 1/2 for the iGPU. The qwen3-coder-30b 8bit quant model was (as you would expect) only taking 30GB (a bit less than half of what was allocated). Though weirdly, in htop it shows that the CPU has 125GB available to it, so I'm not sure what to make of that.
- alias_neo 1y agoI'm pretty new to this, so if I wanted to benchmark my current hardware and compare to your results what would be the best way to do that? I'm looking at going for a Framework Desktop and would like to know what kind of performance gain I'd get over the current hardware I have, which so far I have a "feel" for the performance of from running Ollama and OpenWebUI, but no hard numbers.
- rjzzleep 1y agoMaybe the real value of the DGX spark is to work on Switch 2 emulation. ARM + Nvidia GPU. Start with Switch 2 emulation on this machine and then optimize for others. (Yeah, I know, kind of expensive toy).
- pta2002 1y agoI think you can get something a lot cheaper if that’s all you want, e.g. something in the Jetson Orin line. That’s more similar to the switch, also, since it’s a Tegra CPU.
- ThatMedicIsASpy 1y agoExpensive today. But how quickly (years) will these systems lower in value? At least on the Nvidia side of things they can be stacked.. so maybe not so much =/
- lhl 1y agoRDNA3 CUs do not have FP8 support and its INT8 runs at the same speed as FP16 so Strix Halo's max theoretical is basically 60 TFLOPS no matter how you slice it (well it has double INT4, but I'm unclear on how generally useful that is): 512 ops/clock/CU * 40 CU * 2.9e9 clock / 1e12 = 59.392 FP16 TFLOPS Note, even with all my latest manual compilation whistles and the latest TheRock ROCm builds the best I've gotten mamf-finder up to about 35 TFLOPS, which is still not amazing efficiency (most Nvidia cards are at 70-80%), although a huge improvement over the single-digit TFLOPS you might get ootb. If you're not training, your inference speed will largely be limited by available memory bandwidth, so the Spark token generation will be about the same as the 395. On general utility, I will say that the 16 Zen5 cores are impressive. It beats my 24C EPYC 9274F in single and multithreaded workloads by about 25%.
- littlestymaar 1y agoYou should add memory bandwidth to your comparison, as it's usually the bottleneck in terms of tps (at least for token generation, prompt processing is a different story).
- robbomacrae 1y agoGosuCoder's latest video seems to be a well timed test of using Ryzen AI Max on some local models getting 40 TPS on a quantized Qwen 3 coder. https://www.youtube.com/watch?v=0DET4YFzS6A https://www.youtube.com/watch?v=0DET4YFzS6A