6 ms·
Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio. I have a Strix Halo and dual 32GB GPUs in my desktop, a
by SwellJoe 2mo ago
Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio.
I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context.
And, MoE should make it run at a close to usable speed.
- jubilanti 2mo agoA 3060ti 8gb, released in 2020, has 448 GB/s of bandwidth compared to the Halo 256 GB/s The 3080ti is 912.4 GB/s
- embedding-shape 2mo agoAnd the newly announced/launched Apple M6 has 170GB/s of unified memory bandwidth, meanwhile M5 Ultra gets 1.2TB/s of unified memory bandwidth. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/ https://www.apple.com/newsroom/2026/08/apple-introduces-m6-a... Not sure if the first one is a typo on their press release, can't be just 170GB/s then be pushed for AI use, can it? Could be a different measurement I suppose...
- jtbayly 2mo agoYou got me curious so I looked up the previous chips[0]. Memory bandwidth M1: 68 GB/s M2: 100 GB/s (47% increase) M3: 100 GB/s (0% increase) M4: 120 GB/s (20% increase) M5: 153 GB/s (27.5% increase) So, M6: 170 GB/s (11% increase) doesn’t seem impossible, though I would have expected more. [0]: https://www.jdhodges.com/blog/apple-cpu-compared-m1-m3-m3-m4-m5-max/ https://www.jdhodges.com/blog/apple-cpu-compared-m1-m3-m3-m4...
- mhast 2mo agoThe different models of chips and memory config have very different memory speeds as well. Eg the M4 Max 128GB has a bandwidth speed of 500GB/s+. And that's true for other models as well. But as you note, the base speed has also increased over the versions.
- brandall10 2mo agoFrom a bandwidth perspective, the ultra is like 8 M5s fused together (@ 150GB/s), that's how it gets to the 1200. Historically the Pro doubles the base, the Max doubles the Pro, and the Ultra doubles the Max. If an M6 ultra were released today it would be 1.36TB/s.
- DwarvenEngineer 2mo agodoes that mean they're measuring bandwidth differently than how others (like nvidia) does it? memory bandwidth is the gating factor of running models locally, so if it's actually 8x 150GB/s, it may help something like prefill, but would it actually speed up decode comparatively?
- entrope 2mo agoNo, they use the same definition of memory bandwidth as others, but Apple Silicon has a lot of memory channels. In previous generations, prefill has been compute-limited and decode is fast. https://blog.exolabs.net/nvidia-dgx-spark/ https://blog.exolabs.net/nvidia-dgx-spark/ outlines a combination of a DGX Spark and an M3 Ultra that took advantage of fast prefill on the Nvidia hardware and fast decode on Apple Silicon.
- ActorNightly 2mo agoGFX vram is still faster.
- pizza234 2mo agoNot a hardware engineer, but it's mainly because of RAM wires/channels (not implying that this is "simple" form an engineering perspective). Using the published bandwidths, the math is 170 * 1 and 153 * 8.
- embedding-shape 2mo agoBut 170GB/s is almost nothing? None of the RTX 50 series GPUs has that low bandwidth, you have to go back two generations of nvidia GPUs to get closer to that, and then it's the cheapest of the series, RTX 3050, which has ~170GB/s. Even the GTX 1080, launched ten years ago, has double the bandwidth! This must be some different way of measuring the bandwidth right? Since they explicitly say this for AI, but the numbers they share don't show that at all. Or I gravely misunderstand something here.
- pizza234 1mo ago> Since they explicitly say this for AI That's marketing spin indeed (or lies, if you prefer).
- SwellJoe 2mo agoEnjoy Gemma 4 E2B at blistering speeds, I guess?
- jubilanti 2mo agoMy point was more: this was the 2nd lowest end card from a generation 6 years ago, and it had way higher bandwidth than today's alleged flagship.
- jcsycombinator 2mo agoSmall amounts of fast ram vs huge amounts of slower ram. It costs more than the strix to just buy regular ddr5 ram sticks today.
- yjftsjthsd-h 2mo agoBut it also has 8GB of RAM.
- ActorNightly 2mo ago[flagged]
- sosodev 2mo agoThat’s only true if you think AI is the only reason to own a powerful and efficient server. Mine does plenty of traditional server stuff too.
- ArvidSu 2mo agoAn "AI" server can do traditional server stuff but a traditional server can't do AI stuff (inference)
- SwellJoe 2mo agoI can do traditional server stuff on any old computer with a big hard disk and a decent amount of RAM. That's not worth $3500-$4000. When RAMpocalypse is over and we can buy a Strix Halo for under $2000 again, the math starts mathing. It becomes a pretty great desktop computer that also happens to run AI pretty well at a pretty good price.
- sosodev 2mo agoYeah, but that computer can’t also do the AI stuff. And not everybody has a desktop with multiple 32GB GPUs available. I’ll admit though I’m biased because I bought my board for $1600 back before the prices went crazy.
- downrightmike 2mo agoOh no a tough constraint that will lead to further innovation like deepseek. How terrible.
- pizza234 2mo agoI spent around 5k on a server for "AI stuff" and it's currently doing no AI, because local LLMs (at least on systems with 32 GB VRAM) can only do only very basic stuff; this includes Qwen3.8 - in spite of the reverse engineering blog post, when I've tried Qwen to do a similar task, it flunked miserably. Additionally, I've read on some informal sources, the next step in quality is at 256 GB, not 128, which is very expensive (it's around 10k). 10k for privacy is... a toy for rich tinkerers, considering that most the people have their email on the cloud.
- throwaw12 2mo agohow much performance (tok/s) can you expect from 128GB Strix Halo? assuming this model will be released with FP8 also can you use it for fine tuning?
- SwellJoe 2mo agoThe Strix Halo and DGX Spark are pretty danged slow, relatively speaking. I don't recall exact numbers, but with MoE models in this size ballpark (Laguna S 2.1), I seem to recall I was seeing about 20-25 t/s with a big context, which is close to usable. Qwen 3.8 27B crawls on this hardware, though, at 10-16 t/s, definitely not comfortable for interactive use. (Though this makes it seem like you can cook pretty good with a 4-bit ROCmFP4 quantization: https://github.com/julianmb/q38rocm https://github.com/julianmb/q38rocm the model does get notably dumber below six bits.) A model similar in size to Laguna S 2.1, but with only 6B active parameters, should be a notable amount faster, so I would imagine 25-30 t/s would be a reasonable guess for where Qwen 3.8 Flash Next will land. DFlash2 might improve all these numbers. It wasn't available last I was testing new models on the Strix Halo; I've only used MTP (which doesn't generally improve MoE models, but I believe DFlash2 can). Given software improvements, I'm hopeful an MoE in this size range will be the sweet spot that pushes past 40 t/s and is also smart enough for real work. Qwen 3.8 27B is finally a self-hostable model that's smart enough, but it thinks so hard it still isn't really useful for agentic interactive use. Note also prefill with large models is pretty slow on the Strix Halo (300 t/s, maybe). Time to first token is a painful wait, when using it interactively with large models.
- decide1000 2mo agoOn the DGX I get 44.5 tokens per second (NVFP4). With 8 concurrent it's 241 t/s total. I am using the PrismaAQUA standard 9.7 t/s + Dflash2 30 t/s + torch-compile 37 t/s c8 = 177 t/s
- SwellJoe 2mo agoWhat model? Also, I don't know what "the PrismaAQUA" means, ddg thinks it's a CPAP machine, which seems unlikely to help with inference performance. Also, 4-bit has measurable intelligence loss. Sometimes worth it, but, at this size models are barely smart enough at 8 or 6.
- AbsurdCensor 2mo agoI have had an impossible time getting 120B or better models running on Strix Halo (especially under Windows) with any large context windows. And 30-40 tokens/second is fine, but not the fastest. For the most part lately I have been sticking with Qwen 3.8 27b and that thing will easily suck up 64gb of ram. Add in docker with some additional programs running and it's really easy to eat up 128gb of ram.
- SwellJoe 2mo agoI found a couple of different 4-bit quantizations of Laguna S 2.1 that run pretty well with pretty big context (also quantized, to 8 bits, I think). Unfortunately, Laguna isn't better than Qwen 3.8 27B, which I'm able to run at roughly the same speed on my desktop machine, so I don't use Laguna or the Strix Halo very much, lately. (It's also too hot for me to be running heaters for inference. It's been ~110F most days for the past few weeks.)
- deleted 1mo ago[deleted]