7 ms·
Dell's version of the DGX Spark fixes pain points
- kachapopopow 9mo agoDell fixing issues instead of creating new ones? That's a new one for me. Would rather still not deal with their firmware updaters thought.
- cjbgkagh 9mo agoGive them a chance, I’m sure they’ll add new issues in one of their monthly bios updates.
- kachapopopow 9mo agonothing beats perfectly good vendor firmware updates packaged in an obscenely complicated bash file that just extracts the tool and runs it while performing unnecessary and often broken validation that only runs on hardware that is part of their ecosystem (ex: dell nic on non dell chassis).
- BadBadJellyBean 9mo agoOn linux I use fwupdmgr to upgrade the firmware on my dell laptop. Not sure if that works for servers though.
- buildbot 9mo agoIndeed, same process here: https://www.dell.com/support/kbdoc/en-us/000379162/how-to-upgrade-the-bios-and-the-firmware-on-a-dell-pro-max-with-the-grace-blackwell-system https://www.dell.com/support/kbdoc/en-us/000379162/how-to-up...
- kachapopopow 9mo agotends to be hit or miss when you use dell parts on non dell hardware (but the cost savings are worth it since typically nobody wants to touch dell hardware due to these issues)
- Tepix 9mo agoYou can get two Strix Halo PCs with similar specs for that $4000 price. I just hope that prompt preprocessing speeds will continue to improve, because Strix Halo is still quite slow in that regard. Then there is the networking. While Strix Halo systems come with two USB4 40Gbit/s ports, it's difficult to a) connect more than 3 machines with two ports each b) get more than 23GBit/s or so per connection, if you're lucky. Latency will also be in the 0.2ms range, which leaves room for improvement. Something like Apple's RDMA via Thunderbolt would be great to have on Strix Halo…
- Aurornis 9mo agoThe primary advantage of the DGX box is that it gives you access to the nVidia ecosystem. You can develop against it almost like a mini version of the big servers you're targeting. It's not really intended to be a great value box for running LLMs at home. Jeff Geerling talks about this in the article.
- cmrdporcupine 9mo agoExactly this. I'm not sure why people keep drumming the "a Mac or Strix Halo is faster/cheaper" drum. Different market. If I want to do hobby / amateur AI research or do stuff with fine tuning models etc, learn the tooling. I'm better off with the DG10 than AMD or Apple's systems. The Strix Halo machines look nice. I'd like one of those too. Especially if/when they ever get around to getting it into a compelling laptop. But I ordered the ASUS Ascent DG10 machine (since it was more easily available for me than the other versions of these) because I want to play around with fine tuning open weight models, learning tooling, etc. That and I like the idea of having a (non-Apple) Aarch64 linux workstation at home. Now if the courier would just get their shit together and actually deliver the thing...
- mapontosevenths 9mo agoI have this device, it's exactly as you say. This is a device for AI research and development. My buddies mac ultra beats it squarely for inference workloads, but for real tinkering it can't be beat. I've used it to fine tune 20+ models in the last couple of weeks. Neither a Mac or Strix Halo even try to compete.
- jasoneckert 9mo agoI've got the Dell version of the DGX Spark as well, and was very impressed with the build quality overall. Like Jeff Geerling noted, the fans are super quiet. And since I don't keep it powered on continuously and mainly connect to it remotely, the LED is a nice quick check for power. But the nicest addition Dell made in my opinion is the retro 90's UNIX workstation-style wallpaper: https://jasoneckert.github.io/myblog/grace-blackwell/ https://jasoneckert.github.io/myblog/grace-blackwell/
- ranger_danger 9mo agoI just want a standard, affordable mini PC that looks like this one. Or better yet, with the brown accents normally found on recent PowerEdge systems. https://www.fsi-embedded.jp/contents/uploads/2018/11/DELLEMC_R740_24x25_bezel_2_above_ff.jpg https://www.fsi-embedded.jp/contents/uploads/2018/11/DELLEMC...
- storus 9mo agoZotac has a bunch of x64 mini PCs that use a similar hexagonal styling.
- mapontosevenths 9mo agoI've had mine for a while now, and never actually connected a monitor to it. Now I'll have to. Thanks. :)
- alecco 9mo agoIMHO DGX Spark at $4,000 is a bad deal with only 273 GB/s bandwidth and the compute capacity between a 5070 and a 5070 TI. And with PCIe 5.0 at 64 GB/s it's not such a big difference. And the 2x 200 GBit/s QSFP... why would you stack a bunch of these? Does anybody actually use them in day-to-day work/research? I liked the idea until the final specs came out.
- BadBadJellyBean 9mo agoI think the selling point is the 128GB of unified system memory. With that you can run some interesting models. The 5090 maxes out at 32GB. And they cost about $3000 and more at the moment.
- alecco 9mo ago1. /r/localllama unanimously doesn't like the Spark for running models 2. and for CUDA dev it's not worth the crazy price when you can dev on a cheap RTX and then rent a GH or GB server for a couple of days if you need to adjust compatibility and scaling.
- BadBadJellyBean 9mo agoI am not on reddit. What are they saying?
- mapontosevenths 9mo agoIt isn't for "running models." Inference workloads like that are faster on a mac studio, if that's the goal. Apple has faster memory. These devices are for AI R&D. If you need to build models or fine tune them locally they're great. That said, I run GPT-OSS 120B on mine and it's 'fine'. I spend some time waiting on it, but the fact that I can run such a large model locally at a "reasonable" speed is still kind of impressive to me. It's REALLY fast for diffusion as well. If you're into image/video generation it's kind of awesome. All that compute really shines when for workloads that aren't memory speed bound.
- colordrops 9mo agoI assume they didn't fix the memory bandwidth pain point though.
- llm_nerd 9mo agoThe memory bandwidth limitation is baked into the GB10, and every vendor is going to be very similar there. I'm really curious to see how things shift when the M5 Ultra with "tensor" matmul functionality in the GPU cores rolls out. This should be a multiples speed up of that platform.
- storus 9mo agoMy guess is M5 Ultra will be like DGX Spark for token prefill and M3 Ultra for token generation, i.e. the best of both worlds, at FP4. Right now you can combine Spark with M3U, the former streaming the compute, lowering TTFT, the latter doing the token generation part; with M5U that should no longer be necessary. However given RAM prices situation I am wondering if M5U will ever get close to the price/performance of Spark + M3U we have right now.
- echion 9mo ago> you can combine Spark with M3U, the former streaming the compute, lowering TTFT, the latter doing the token generation part Are you doing this with vLLM, or some other model-running library/setup?
- coder543 9mo agoThey're probably referencing this article: https://blog.exolabs.net/nvidia-dgx-spark/ https://blog.exolabs.net/nvidia-dgx-spark/
- kristianp 9mo agoThe M3 ultra was released about 18 months after the original M3, so you could be waiting a while for the M5 Ultra.
- dagaci 9mo agoA nice little AI review with comparison of the CPU/Power Draw & Networking would be interested in seeing a fine-tuning comparison too. I think pricing was missing also.
- geerlingguy 9mo agoI've been working on fine tuning testing, it's something I hope to set up for comparison against the Mac Studio and Framework Desktop clusters soon.
- npalli 9mo agoSeems you are paying the Dell tax of 15%. The same setup is $4K from NVidia, Lenovo and $3K for 1TB at Asus. https://www.dell.com/en-us/shop/desktop-computers/dell-pro-max-with-gb10/spd/dell-pro-max-fcm1253-micro/xcto_fcm1253_usx https://www.dell.com/en-us/shop/desktop-computers/dell-pro-m...
- kristianp 9mo agoI know it's just a quick test, but llama 3.1 is getting a bit old. I would have liked to see a newer model that can fit, such as gpt-oss-120, (gpt-oss-120b-mxfp4.gguf), which is about 60gb of weights (1). (1) https://github.com/ggml-org/llama.cpp/discussions/15396 https://github.com/ggml-org/llama.cpp/discussions/15396
- eurekin 9mo agoCorrect, most of r/LocalLlama moved onto next gen MoE models mostly. Deepseek introduced few good optimizations that every new model seems to use now too. Llama 4 was generally seen as a fiasco and Meta haven't made a release since
- fragmede 9mo agoWhat are some of the models people are using? (Rather than naming the ones they aren't.)
- eurekin 9mo agoGLM 4.7 is new and promising. MinMax 2.1 is good for agents. Of course the qwen3 family, vl versions are spectacular. NVIDIA Nemotron Nano 3 excels at long context and the unsloth variant has been extended to 1m tokens. I thought the last one was a toy, until I tried with a full 1.2 megabyte repomix project dump. It actually works quite well for general code comprehension across the whole codebase, CI scripts included. Gpt-oss-120 is good too, altough I'm yet to try it out for coding specifically
- nightski 9mo agoDoes GLM 4.7 run well on the spark? I thought I read it didn’t but it wasn’t clear.
- magicalhippo 9mo agoSince I'm just a pleb with a 5090, I run GPT-OSS 20B a lot, since it fits comfortably in VRAM with max context size. I find it quite decent for a lot of things, especially after I set reasoning effort to high and disabled top-k and top-p and set min-p to something like 0.05. For the Qwen3-VL, I recently read that someone got significantly better results by using F16 or even F32 versions of the vision model part, while using a Q4 or similar for the text model part. In llama.cpp you can specify these separately[1]. Since the vision model part is usually quite small in comparison, this isn't as rough as it sounds. Haven't had a chance to test that yet though. [1]: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md https://github.com/ggml-org/llama.cpp/blob/master/tools/serv... (using --mmproj AFAIK)
- barelysapient 9mo agoGreat article but would be nice to see how larger models work.
- geerlingguy 9mo agoSee: https://github.com/geerlingguy/ai-benchmarks/issues/34 https://github.com/geerlingguy/ai-benchmarks/issues/34
- cat_plus_plus 9mo agoI have a slightly cheaper similar box, NVIDIA Thor Dev Kit. The point is exactly to avoid deploying code to servers that cost half a million dollars each. It's quite capable in running or training smart LLMs like Qwen3-Next-80B-A3B-Instruct-NVFP4. So long as you don't tear your hair out first figuring out pecularities and fighting with bleeding edge nightly vLLM builds.
- echion 9mo ago> training smart LLMs like Qwen3-Next-80B-A3B-Instruct-NVFP4 Sounds interesting; can you suggest any good discussions of this (on the web)?
- nightski 9mo agoIt's a product without a purpose.
- postalrat 9mo agoSpark's biggest paint point is the price. Does it fix that?
- bigyabai 9mo agoThere's an entire line of Linux-supported Jetson products available for your perusal, in addition to all of the GTX and RTX cards that have native ARM64 support.
- mmaunder 9mo agoFor those of you wondering if this fits your use case vs the RTX 5090 the short answer is this: The desktop RTX 5090 has 1792 GB/s of memory bandwidth partially due to the 512 bit bus width, compared to the DGX Spark with a 256 bit bus and 273 GB/s memory bandwidth. The RTX 5090 has 32G of VRAM vs the 128G of “VRAM” in the DGX Spark which is really unified memory. Also the RTX 5090 has 21760 cuda cores vs 6144 in the DGX Spark. (3.5 x as many). And with the much higher bandwidth in the 5090 you have a better shot at keeping them fed. So for embarrassingly parallel workloads the 5090 crushes the Spark. So if you need to fit big models into VRAM and don’t care about speed too much because you are for example, building something on your desktop that’ll run on data center hardware in production, the DGX Spark is your answer. If you need speed and 32G of VRAM is plenty, and you don’t care about modeling network interconnections in production, then the RTX 5090 is what you want.
- chao- 9mo agoIt's also worth nothing that the 128GB of "VRAM" in the GB10 is even less straightforward than just being aware that the memory is shared with the CPU cores. There's a lot of details in memory performance that differ across both the different core types, and the two core clusters: https://chipsandcheese.com/p/inside-nvidia-gb10s-memory-subsystem https://chipsandcheese.com/p/inside-nvidia-gb10s-memory-subs...
- kouteiheika 9mo ago> building something on your desktop that’ll run on data center hardware in production, the DGX Spark is your answer It isn't, because it's a different architecture than the datacenter hardware. They're both called "Blackwell", but that's a lie[1] and you still need "real" datacenter Blackwell card for development work. (For example, you can't configure/tune vLLM on Spark, and then move it into a B200 and even expect it to work, etc.) [1] -- https://github.com/NVIDIA/dgx-spark-playbooks/issues/22 https://github.com/NVIDIA/dgx-spark-playbooks/issues/22
- benreesman 9mo agosm_120 (aka 1CTA) supports tensor cores and TMEM just fine: example 83 shows block-scaled NVFP4 (I've gotten 1850 ish dense TFLOPs at 600W, the 300W part caps out more like 1150). sage3 (which is no way in hell from China, myelin knows it by heart) cracks a petaflop in bidirectional noncausal. The nvfuser code doesn't even call it sm_100 vs. sm_120: NVIDIA's internal nomenclature seems to be 2CTA/1CTA, it's a bin. So there are less MMA tilings in the released ISA as of 13.1 / r85 44. The mnemonic tcgen05.mma doesn't mean anything, it's lowered onto real SASS. FWIW the people I know doing their own drivers say the whole ISA is there, but it doesn't matter. The family of mnemonics that hits the "Jensen Keynote" path is roughly here: https://docs.nvidia.com/cuda/parallel-thread-execution/#warp-level-block-scaling https://docs.nvidia.com/cuda/parallel-thread-execution/#warp.... 10x path is hot today on Thor, Spark, 5090, 6000, and data center. Getting it to trigger reliably on real tilings? Well that's the game just now. :) Edit: https://customer-1qh1li9jygphkssl.cloudflarestream.com/1795aaa9e141e3b87546368033bf77ef/watch https://customer-1qh1li9jygphkssl.cloudflarestream.com/1795a...
- cpgxiii 9mo agoAbsent disassembly and direct comparison between a DGX Spark and a Dell GB10, I don't think there's sufficient evidence to say what is meaningfully different between these devices (beyond the obvious of the power LED). Anything over 240W is beyond the USB-C EPR spec, and while Dell does have a question ably-compliant USB-C 280W supply, you'd have to compare actual power consumption to see if the Dell supply is actually providing more power. I suspect any other minor differences in experience/performance are more explainable as the consequences on increasing maturity of the DGX software stack than anything unique to the Dell version; particularly any comparisons to very early DGX Spark behavior need to keep in mind that the software and firmware have seen a number of updates.
- geerlingguy 9mo agoComparing notes with Wendell from Level1Techs, the ASUS and Dell GB10 boxes were both able to sustain better performance due to their better thermal management. That's a fairly significant improvement. The Spark's crusted gold facade seems more form over function.
- graham33 9mo agoI have NixOS running on my DGX Spark: https://github.com/graham33/nixos-dgx-spark https://github.com/graham33/nixos-dgx-spark, would be interested to know if the USB image also boots on the Dell Pro Max GB10.
- supermatt 9mo agoJeff, This is the second time you have been given a prosumer level cluster pretty much built for local LLM inference and on both occasions you have performed benchmarks without batching. If you still have the hardware (this and the Mac cluster) can you PLEASE get some advice and run some actually useful benchmarks? Batching on a single consumer GPU often results in 3-4x the throughput. We have literally no idea what that batching looks like on a $10k+ cluster without otherwise dropping the cash to find out.