3 ms·
Please don’t buy a DGX Spark unless all three of these are true: - You value simplicity more than performance or price-to-performance. - You accept that th
by biddit 2mo ago
Please don’t buy a DGX Spark unless all three of these are true:
- You value simplicity more than performance or price-to-performance.
- You accept that the hardware will depreciate rapidly.
- You’re prepared to buy two or four of them.
OR:
- You want to run frontier models right now as cheaply as possible
- You want to run high-parameter models on a 15a breaker/line
Otherwise, get a normal, high-bandwidth GPU.
A single Spark gives you roughly 115 GB of usable memory compared with the 24–32 GB found on many lower-cost GPUs. It's certainly a big increase, but in practice it does not unlock dramatically better models.
- One Spark: More memory, but mostly enough for poor-quality, extremely low-bit quants of larger models.
- Two Sparks: Enough for mid-tier parameter models at reasonable quants, such as DeepSeek V4 Flash and HY3.
- Four Sparks: Enough for GLM 5.2 at a reasonable quant. You'll need a $1000+ switch too.
The problem is that Sparks are slow compared with almost everything else in their price range. Many factors affect inference speed, but memory bandwidth is one of the biggest. A $4,000-plus DGX Spark provides only 273 GB/s.
| GPU | Memory bandwidth | VRAM | Approx. price |
| ------------------- | ---------------: | -------------: | ------------: |
| DGX Spark | 273 GB/s | ~115 GB usable | $4,000+ |
| RTX 5060 | 448 GB/s | 16 GB | $600 |
| Radeon AI Pro R9700 | 640 GB/s | 32 GB | $1,200 |
| RTX 4000 Pro | 672 GB/s | 24 GB | $2,300 |
| RTX 4500 Pro | 896 GB/s | 32 GB | $3,500 |
| RTX 3090 | 936 GB/s | 24 GB | $1,200 |
| RTX 5000 Pro | 1,344 GB/s | 48 GB | $6,000 |
| RTX 5090 | 1,792 GB/s | 32 GB | $4,000 |
| RTX 6000 Pro | 1,792 GB/s | 96 GB | $12,000 |
Yes, the Spark has substantially more memory. But going from roughly 24 GB to 115 GB does not necessarily unlock substantially better model quality. In many cases, it only lets you load heavily compressed 2-bit versions of larger models, such as DeepSeek V4 Flash, with serious quality degradation.
24–32 GB is currently a sweet spot. Models such as Qwen 3.6 27B and 35B-A3B:
- Perform far above what their parameter counts suggest.
- Fit comfortably within 24–32 GB of VRAM at reasonable quantization levels.
A 4-bit quant of Qwen 3.6 27b (18 GB) will out-perform a 2-bit quant of DeepSeek v4 Flash (90gb).
Instead of the Spark, if I had a roughly $4,000 budget...
Assuming I already had a reasonably modern desktop:
- One RTX 5090, RTX 5000 Pro, or RTX 4500 Pro.
- Two RTX 3090s, RTX 4000 Pros, or R9700s, provided the motherboard can bifurcate two physical x16 slots into x8/x8.
If I were building a system from scratch:
- A DDR4- or PCIe 4.0-era consumer CPU and motherboard that supports x8/x8 bifurcation.
- Two RTX 3090s, RTX 4000 Pros, or R9700s.
If I were already planning to buy a new Mac:
- A MacBook Pro M5 with 64 GB or 128 GB of unified memory.
For context, these are the systems I currently run:
- EPYC Turin with four RTX 6000 Pro Max-Qs.
- EPYC Milan with four RTX 3090s.
- AM4 with two RTX 3090s.
- AM4 with two RTX 3090s.
- Intel Raptor Lake with two RTX 5060 Ti.
- MacBook Pro M3 128GB Unified
- pizza234 2mo agoI'm actually confused by the post - indeed why choosing the 27B dense and the 35B MoE models? Given the structure of the post I honestly have the impression that the Spark was purchased out of curiosity more than for a specific purpose (LLMs).
- biddit 2mo agoAh yeah, your reaction makes sense. In the circles I run in, there is a lot of hype around Sparks for inference, so my gut reaction is to respond with this type of warning. I did not intend to imply that the post author was advocating that they're great for inference, as they're obviously not.
- InTheArena 2mo agoThe 27b dense is _remarkably_ better at coding tasks. Especially at larger quanitizations (Q4 is pretty crap).
- gerdesj 2mo agoActually, its ~119GB usable (I have one). You can shutdown a lot of unneeded services if you only use it remote which will trim a lot more fat. You can enable the RDP service if you don't want to sit in front of it but get a desktop interface. The "shitty" network is 10Gb/s and wifi7! You get a twin QSFP28DDlol+++ (I jest) that each run at 200Gb/s - not for the casual home user but handy at work, although I "only" have 40Gb/s on my switches sigh. With and no switch two you can do a three node cluster with some careful networking. If you want to do more then a switch is needed and it will need to be pretty funky! That said you could wire them up in a circle and use VLANs and MSTP and accept less than 200Gb/s per link. You'll probably need Openvswitch and a lie down afterwards. I'm not a fan of the Gnome desktop but it works well enough and I think the Nvidia customised Ubuntu is well thought out. You get all the complicated NVidia extras pre-installed, along with docker (full fat, not the Ubuntu one) for a fairly quick start. It includes Ubuntu Pro which is free for five systems anyway but its nice to see it pre-installed. You can run quite decent models on this thing see: https://spark-arena.com/ https://spark-arena.com/ Also see "DS4". We blew abut £4000 on one and it will pay for itself in a few months. I tried pricing up an Apple thingie and the Store wouldn't offer me more than 96Gb of RAM and a delivery date in Q3 at the earliest. Our Spark rocked up next day. They seem to come in 1TB or 4TB SSD variants. 1TB is enough for me and saves a lot of cash - keep an eye on your model downloads and ruthlessly delete old experiments. docker system prune. We went for the Asus variant that has active cooling and I stuck it in the ceiling cable tray over our computer room racks. It sits on 1½" stainless steel mesh with lots of clearance in an actively cooled environment.