10 ms·
Nvidia DGX Spark: When benchmark numbers meet production reality
- RyeCatcher 11mo agoWould love to hear from others using the spark for model training and development.
- stuckinhell 11mo agoI'm utterly shocked at the article saying GPU inference (PyTorch/Transformers)isn't working. Numerical instability produces bad outputs, Not viable for real-time serving, Wait for driver/CUDA updates! My job just got me and our entire team a DGX spark. I'm impressed at the ease of use for ollama models I couldn't run on my laptop. gpt-oss:120b is shockingly better than what I thought it would be from running the 20b model on my laptop. The DGX has changed my mind about the future being small specialized models.
- jasonjmcghee 11mo ago> I'm utterly shocked at the article saying GPU inference (PyTorch/Transformers)isn't working Are you shocked because that isn't your experience? From the article it sounds like ollama runs cpu inference not GPU inference. Is that the case for you?
- RyeCatcher 11mo agoTotally agree. I’ve been training nanochat models all morning. Hit some speed bumps. I’ll share more later in another article. Buts it’s absolutely amazing. I fine tuned a Gemma3 model in a day yesterday.
- deleted 11mo ago[deleted]
- deleted 11mo ago[deleted]
- jsheard 11mo agoNo mention of the monstrous 200GbE NIC, seems like a waste if people aren't finding a use for it.
- RyeCatcher 11mo agoNeed to buy 2 and connect em. :-)
- RyeCatcher 11mo agoI absolutely love it. I’ve been up for days playing with it. But there are some bleeding edge issues. I tried to write a balanced article. I would highly recommend for people that love to get their hands dirty. Blows away any consumer GPU.
- furyofantares 11mo agoSince the text is obviously LLM output, how much prompting and editing went into this post? Did you have to correct anything that you put into it that it then got wrong or added incorrect output to?
- NathanielK 11mo agoDefinitely reeks of someone who doesn't know what makes a readable blogpost and hoped the LLM did. I was not familiar with the hardware, so I was disappointed there wasn't a picture of the device. Tried to skim the article and it's a mess. Inconsistent formatting and emoji without a single graph to visualize benchmarks.
- furyofantares 11mo agoI read the whole thing now and it's filled with slop. I don't really care about the emojis and the marketing voice too much. I do care that it's impossible to tell what the author cared about what they didn't, or if any of it is made up or extrapolated. I bet the input to the LLM would have been more interesting.
- furyofantares 11mo ago> Training Performance is Real (When It Works) It looks like it worked? Why's it say this? > Verdict: Inference speed scales proportionally with model size. Author only tried one model size and it's faster than NVIDIA's reported speed at a larger model. Not really a "Verdict". > Verdict: 4-bit quantization is production-viable. That's not really something you can conclude from messing around with it and saying you like the outputs. > GPU Inference is Fundamentally Broken Probably not? It probably just doesn't work in llama.cpp right now? Takes a while reading this to work out they tried ollama and then later llama.cpp, which I'd guess is basically testing llama.cpp twice. Actually I don't even believe that, I'm sure author ran into errors that might be a pain to figure out, but there's no evidence it's worse than that. But then it says this is the "root cause": ARM64 + Blackwell + CUDA 13.0 = Bleeding Edge ↓ Limited production testing ↓ Edge cases in numerical precision (inference) ↓ Memory management issues (training) Am I to believe GPU inference is really fundamentally broken? I'm not seeing the case made here, just claims. At this point the LLM seems to have gotten confused about whether it's talking about the memory fragmentation issue or the GPU inference issue. But it's hard to believe anything from this point on in the post.
- veber-alex 11mo agoThe llama.cpp issues are strange. There are official benchmarks of the Spark running multiple models just fine on llama.cpp https://github.com/ggml-org/llama.cpp/discussions/16578 https://github.com/ggml-org/llama.cpp/discussions/16578
- RyeCatcher 11mo agoCool I’ll have a look. All reflections I made were first pass stuff.
- CaptainOfCoit 11mo agoThere wasn't any instructions how the author got ollama/llama.cpp, could possibly be something nvidia shipped with the DGX Spark and is an old version?
- moffkalast 11mo agoLlama.cpp main branch doesn't run on Orins so it's actually weird that it does run on the Spark.
- eadwu 11mo agoThere are bleeding edge issues, everyone dials into transformers so that's generally pain proof. I haven't exactly bisected the issue but I'm pretty sure convolutions are broken on sm_121 after a certain size, getting 20x memory blowup from a convolution from a 2x batch size increase _only_ on the DGX Spark. I haven't had any problems with inference, but I also don't use the transformers library that much. llama.cpp was working for openai-oss last time I checked and on release, not sure if something broke along the way. I don't exactly know if memory fragmentation is something fixable on the driver side - this might just be the problem with kernel's policy and GPL, it prevents them from automatically interfering with the memory subsystem to the granularity they'd like - see zfs and their page table antics - or so my thoughts on it is. If you've done stuff on WSL, you have similar issues and you can fix it by running a service that normally compacts and clean memory, I have it run every hour. Note that this does impact at the very least CPU performance and memory allocation speeds, but I have not have any issue with long training runs with it (24hr+, assuming that is the issue, I have never tried without it and put that service in place since getting it due to my experience on WSL).
- suprjami 11mo agoSo I can spend thousands of dollars to have an unstable training environment and inference performance worse than a US$200 3060. Wow. Where do I sign up?
- thehamkercat 11mo agoThe memory bandwidth on this thing is absolute trash, better buy a mac mini/studio with this much ram if you're throwing this much money, it'll be faster (M4 Max)
- suprjami 11mo agoAgree, any Max or Ultra should walk all over this thing, and has the advantage of many years of already-working software. Apple benchmarks: https://github.com/ggml-org/llama.cpp/discussions/4167 https://github.com/ggml-org/llama.cpp/discussions/4167
- bigyabai 11mo agoIt really depends, the metrics are kinda all over the place right now: https://docs.google.com/spreadsheets/d/1SF1u0J2vJ-ou-R_Ry1JZQ0iscOZL8UKHpdVFr85tNLU/edit?gid=0#gid=0 https://docs.google.com/spreadsheets/d/1SF1u0J2vJ-ou-R_Ry1JZ... (cited from https://lmsys.org/blog/2025-10-13-nvidia-dgx-spark/ https://lmsys.org/blog/2025-10-13-nvidia-dgx-spark/)
- vardump 11mo ago3060 doesn't have 128 GB RAM.
- moffkalast 11mo ago128GB / 12 GB = ~11, * 200€ = only 2200€ plus mining rig mobo. It would be cheaper to buy up a dozen 3060s and build a custom PC around them than to buy the Spark.
- 11mo ago
- MaKey 11mo agoWhy would you get this when a Ryzen AI Max+ 395 with 128 GB is a fraction of the price?
- d3m0t3p 11mo agoBecause the ML ecosystem is more mature on the NVidia side. Software-wise the cuda platform is more advanced. It will be hard for AMD to catch up. It is good to see competition tho.
- shikon7 11mo agoBut the article shows that the Nvidia ecosystem isn't that mature either on the DGX Spark with ARM64. I wonder if Nvidia is still ahead for such use cases, all things considered.
- bigyabai 11mo agoOn the DGX Spark, yes. On ARM64, Nvidia has been shipping drivers for years now. The rest of the Linux ecosystem is going to be the problem, most distros and projects don't have anywhere near the incentive Nvidia does to treat ARM like a first-class citizen.
- pjmlp 11mo agoComplete computer with everything working.
- aseipp 11mo agoI'm not yet using mine for ML stuff because there are still a lot of various issues like this post outlined. But I am using mine as an ARM dev system in the meantime, and as a "workstation" it's actually quite good. The Cortex-X925 cores are Zen5 class in performance and it is overall an absolute unit for its size, I'm very impressed that a standard ARM core is pushing this level of performance for a desktop-class machine. I thought about buying a new Linux desktop recently, and this is good enough I might just plug it into a monitor and use it instead. It is also a standard UEFI+ACPI system; one Reddit user even reported that they were able to boot up Fedora 42 and install the open kernel modules no problem. The overall delta/number of specific patches for the Canonical 6.17-nvidia tree is pretty small when I looked (the current kernel is 6.11). That and the likelihood the consumer variant will support Windows hopefully bodes well for its upstream Linux compatibility, I hope. To be fair, most of this also true of Strix Halo from what I can tell (most benchmarks put the DGX furthest ahead at prompt processing and a bit ahead at raw token output. But the software is still buggy and Blackwell is still a bumpy ride overall, so it might get better). But I think it's mostly the pricing that is holding it back. I'm curious what the consumer variant will be priced at.
- eitally 11mo agoOne of my colleagues wrote a first impressions blog post last week. It's from our company's perspective, but is a solid overview of the product and intended capabilities, from the POV of an AI developer or data scientist. https://www.anaconda.com/blog/python-nvidia-dgx-spark-first-impressions https://www.anaconda.com/blog/python-nvidia-dgx-spark-first-...
- victor106 11mo ago< The CPU memory is the same as the GPU memory and is much larger than any other discrete GPU available in a desktop. That means much larger datasets and bigger models can be run locally than would be possible otherwise. Isin't this the same architecture that the Mx from Apple implements from a memory perspective?
- LtdJorge 11mo agoYep, it is
- NathanielK 11mo agoThis is a much better introduction to the hardware.
- CaptainOfCoit 11mo ago> There you’ll see the 10 Cortex-X925 (“performance”) cores listed with a peak clock rate of 4 GHz, along with the 10 Cortex-A725 (“efficiency”) cores listed with a peak clock rate of 2.8 GHz > If you start Python and ask it how many CPU cores you have, it will count both kinds of cores and report 20 > Note that because of the speed difference between the cores, you will want to ensure there is some form of dynamic scheduling in your application that can load balance between the different core types. Sounds like a new type of hell where I now not only need to manage the threads themselves, but also take into account what type of core they run on, and Python straight up report them as the same.
- sidewndr46 11mo ago
- MomsAVoxell 11mo agoSo, it seems like this makes the DGX a viable ARM-based workstation, for those of us who need/want such a thing, while also offering a relatively decent AI/ML environment. Two things need to happen for me to get excited about this: 1. It stimulates other manufacturers into building their own DGX-class workstations. 2. This all eventually gets shipped in a decent laptop product. As much as it pains me, until that happens, it still seems like Apple Sillicon is the more viable option, if not the most ethical.
- gjsman-1000 11mo agoNVIDIA, ethical?
- bigyabai 11mo agoMy heart goes out to all the gamers who discovered they were chopped liver during the crypto boom. Besides that though, I don't see how Nvidia is particularly non-ethical. They cooperate with Khronos, provide high-quality Linux and BSD drivers free of charge, and don't deliberately block third parties from writing drivers to support new standards. From a relativist standpoint that's as sanctimonious as server hardware gets.
- cramsession 11mo agoThey make significant investments in Israel and even said they’d build a new factory there. It doesn’t get any less ethical than that!
- bigyabai 11mo agoAmerican tech leaders often have no other choice. In most states you can be sued for boycotting, divesting or sanctioning Israel for any reason. If you acquire a company with outstanding obligations to Israel, your only option is to fulfill them. Specifically WRT Mellanox, Nvidia's behavior was more petty than callous.
- semessier 11mo agoNvidia products including from the GPU/CUDA libraries world, the NICs and switches tend to feel like MVP frequently. It works in some cases, hopefully in the end but they are far from polished products without rough edges.
- pertymcpert 11mo agoThis article is AI garbage: ARM64 Architecture: Not x86_64 (limited ML ecosystem maturity) No PyTorch wheels for ARM64+CUDA (must use Docker) Most ML tools optimized for x86 No evidence for any of this whatsoever. The author just asked Claude/claude code to write their article and it just plain hallucinated some rubbish.
- bradfa 11mo agoAarch64 and CUDA has been a thing for many years on Jetson boards. Claiming CUDA is immature on arm is very strange.
- furyofantares 11mo agoWe're getting slopped every day now and upvoting it.
- blurbleblurble 11mo agoWe're like little pigs! Like in Upstream Color: https://www.youtube.com/watch?v=zfDyEr8Ykcg https://www.youtube.com/watch?v=zfDyEr8Ykcg
- mgdev 11mo agoYes. Obvious to anyone who writes AI garbage all day.
- enum 11mo ago- https://publish.obsidian.md/aixplore/Practical+Applications/dgx-lab-benchmarks-vs-reality-day-4#%20GPU%20Inference%20is%20Fundamentally%20Broken https://publish.obsidian.md/aixplore/Practical+Applications/... Does it work if you change to torch.bfloat16? - https://publish.obsidian.md/aixplore/Practical+Applications/dgx-lab-benchmarks-vs-reality-day-4#1.+ARM64+Architecture https://publish.obsidian.md/aixplore/Practical+Applications/... The PyTorch 2.9 wheels do work. You can pip install torch --index-url <whatever-it-is> and it just works. You do need to build flash attention from source, which takes an hour or so.
- amelius 11mo agoKind of weird that (gpu) training works but inference doesn't ...
- mgdev 11mo agoThat makes zero sense.
- renaudr 11mo agoHave you tried to run GPT-OSS-120b using TRT-LLM (as you hint NVIDIA probably did it for their benchmark)? https://cookbook.openai.com/articles/gpt-oss/run-nvidia https://cookbook.openai.com/articles/gpt-oss/run-nvidia
- EnPissant 11mo agoDid Nvidia release benchmark numbers for this?
- fxtentacle 11mo ago„273 GB/sec memory bandwidth“ Really? Less RAM bw than an Epyc CPU? And 4x to 8x less than a consumer GPU? How come this doesn’t massively limit LLM inference speeds?
- qskousen 11mo agoIt does - the inference speed is much slower than a consumer video card. The draw for the Spark and systems like it are the massive amounts of memory available to the GPU.
- RyeCatcher 11mo agoAuthor here. I've updated the article based on your feedback. Thank you. Key corrections: Ollama GPU usage - I was wrong. It IS using GPU (verified 96% utilization). My "CPU-optimized backend" claim was incorrect. FP16 vs BF16 - enum caught the critical gap: I trained with BF16, tested inference with FP16 (broken), but never tested BF16 inference. "GPU inference fundamentally broken" was overclaimed. Should be "FP16 has issues, BF16 untested (likely works)." llama.cpp - veber-alex's official benchmark link proves it works. My issues were likely version-specific, not representative. ARM64+CUDA maturity - bradfa was right about Jetson history. ARM64+CUDA is mature. The new combination is Blackwell+ARM64, not ARM64+CUDA itself. The HN community caught my incomplete testing, overclaimed conclusions, and factual errors. Ship early, iterate publicly, accept criticism gracefully. Thanks especially to enum, veber-alex, bradfa, furyofantares, stuckinhell, jasonjmcghee, eadwu, and renaudr. The article is significantly better now.
- colechristensen 11mo agoThis looks like better peer review than most of what gets done for scientific papers.
- anticensor 11mo agoThis is what I would call the value of post-publication review. Pre-publication review is not enough.
- loufe 11mo agoYeah, kudos, OP. It's a very different read before-after.
- Tiberium 11mo agoIs there a reason why you used an LLM for the entire article, and moreover, even for this comment? Couldn't you have at least written this comment yourself?
- CamperBob2 11mo ago
- buyucu 11mo agoStrix Halo from AMD appears to be a much more consumer-friendly alternative than DGX Spark.
- spwa4 11mo agoAm I reading this right? I was expecting much more performance. My 64G M1 Max has 40.72 tok/s on ollama/GPT-OSS-20B (less than half the price of this machine), and M4 Max 128G from a colleague (but 32G would work) gets about 67 tok/s on ollama/GPT-OSS-20B, and apparently the most recent software updates push that to 78 tok/s. The DGX Spark gets 82.74 tok/s. Ryzen Max 395+ gets you 55 tok/s [1] [1] https://www.reddit.com/r/LocalLLaMA/comments/1nabcek/anyone_actully_try_to_run_gptoss120b_or_20b_on_a/ https://www.reddit.com/r/LocalLLaMA/comments/1nabcek/anyone_...