4 ms·
Hands-On with the AMD Ryzen AI Halo
- cmrdporcupine 3mo agoUntil RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (512GB could do a passable job), then I think the story changes. Whether we ever see DRAM as cheap as it was ever again, I don't know.
- science4sail 3mo agoAgree - the 128GB Strix Halo is capable if you use LLMs as assistants, but it's not so good if you use LLMs as agents (or worse, agent teams/swarms) since all of the models that can fit on it are pretty dumb compared to frontier or near-frontier models. You can at best hope for Sonnet-level capabilities. That doesn't mean that local models are useless though! If Mythos/Sol is an ASI that threatens to take your job and turn you into paperclips, then Qwen/Gemma is an old-fashioned office secretary that loyally helps you with tasks but doesn't have a good grasp of details. Every white-collar worker 50 years ago would have killed to have a hard-working personal secretary.
- AmVess 3mo agoExactly this. I own the Framework desktop board. I knew all of its limitations before I bought it, and it's ok to play with on a hobby level, but it isn't much more than a Radio Shack toy. That memory bandwidth is painful. It's like trying to fill an Olympic swimming pool with a thimble. It is excellent as a regular PC, though.
- whaleofatw2022 3mo agoPart of me wonders, would 3d xpoint (if still around) be a viable option? Yeah it is slower than real RAM by a good amount for latency, but you can get similar bandwidth and the cost was history about half of the same size DDR.
- science4sail 3mo agoPossibly - I've heard anecdotal reports that old Intel Optane chips are in hot demand right now. Intel/Micron probably would have made a killing if they had kept that product line alive for a few more years. Never miss an attempt to snatch defeat from the jaws of victory!
- cmrdporcupine 3mo agoI was thought experimenting the other day... ~10 nVME drives striped and running parallel could approximate the memory bandwidth of DDR5 DRAM in a box like this. Like you say, latency wouldn't compete but on raw throughput would be comparable. Not anymore cost effective, I guess, but gets you the ability to work over very large model sizes maybe. But the problem is that tensor matmul etc hardware wouldn't work effectively with it. Useful for KVCache though.
- ciupicri 3mo agoI'll just leave this here: "Achieving 11M IOPS and 66 GB/S IO on a Single ThreadRipper Workstation" (2021), https://news.ycombinator.com/item?id=25956670 https://news.ycombinator.com/item?id=25956670 / https://tanelpoder.com/posts/11m-iops-with-10-ssds-on-amd-threadripper-pro-workstation/ https://tanelpoder.com/posts/11m-iops-with-10-ssds-on-amd-th...
- netinstructions 3mo agoUnfortunately, the table of models and tokens per second (TPS) and time to first token (TTFT) is not helpful without specifying the quantization of the model.
- chmod775 3mo agoThis is more of an ad, not a review, and reads like the author has hardly any experience with the things he's trying out. That Z Image Turbo diffusion model would've also run on many consumer GPUs and with way higher performance for a fraction of the price. Misleading.
- wmf 3mo agoIt's Microcenter. Don't get your reviews from a retailer.
- blitzar 3mo agoSounds like a standard tech reviewer / influencer.
- nrub 3mo agoThat's true, but that's just one model that has some easier setup in the ecosystem. It's not easy to get a consumer gpu that has 128Gb of vram, so while the Halo chip isn't as performant, the ceiling for running larger models is higher.
- syntaxing 3mo agoHighly recommend lemonade server if you have a strix halo desktop. Been using Qwen3.6-35B @ Q_8 as my main driver and it’s been great with 60 TPS for generation. I occasionally use the 27B @ q6 but only get 20-25 TPS for generation with MTP.
- data-ottawa 3mo agoI second this. I used llama swap for a while before leaning into lemonade. The UI has improved a lot, but be careful as the most of the models default to very small 4K context windows by default. They’re doing some nice things with their Halo models, which load an ensemble of different types of model at the same time. With high vram it’s easy to keep them all in memory, so even though the compute is limited the context switching is fast. You do lag a bit on the upstream engine releases, the llama.cpp/sd.cpp/whisper libraries are downloaded from inside the app. vLLM is in experimental mode, I haven’t tested it. It’s limited in the models they suggest, but you can download anything from huggingface with a two click install.
- rkagerer 3mo agoDo you give it access to the internet? Is there some means to monitor the queries it's sending (or hold and review before transmission) or throttle to avoid triggering abuse thresholds on any single domain?
- data-ottawa 3mo agoLemonade proxies your request to llama.cpp that it installs and manages. Is out of the box all local LLMs, but you can connect it to an API endpoint (supports OpenAI compatible or Anthropic compatible). I think it stores basic logs but if you wanted to monitor you’re probably going to want to proxy it with an LLM gateway.
- bobsync 3mo agoseems like an absolute amateur wrote this article
- m463 3mo agoIt is on the microcenter website, so maybe it is intentionally ... mild?
- daft_pink 3mo agoDoes anyone else feel like it would be great to be able to purchase $4000 AI box but 128 gigs is not enough. If I spend all that money and it doesn’t really do what I wanted to do, whats the point? It’s kind of like general aviation where you can go buy a Cessna but it’s only going to realistically get you somewhere you could drive anyways but do you really wanna spend that mush cash to get road trip distance at slightly better than road trip speeds? You really need a 5 million dollar jet and that’s just not practical. That’s sort of how I feel about this device.
- prima-facie 3mo agoThe Strix Halo is a great dev machine and a mediocre AI machine. You can run Qwen 3.6 27B at a decent speed, or larger MoE models, and that's about it. For some that's more than enough though, myself included.
- syntaxing 3mo agoI own one, I don’t feel like the RAM is a huge issue (of course I want 192GB to run something like DS4 Flash). The lack of FP4 and slow memory bandwidth is rough. NVFP4 support is such a huge advantage that I would recommend others to buy a DGX spark over a strix halo if you’re using it purely for AI. Strix halo works better for general computing.
- prima-facie 3mo agoThere's a glimmer of hope with ROCmFP4 which seems to double the current throughput: https://github.com/charlie12345/rocmfp4-llama https://github.com/charlie12345/rocmfp4-llama
- syntaxing 3mo agoI saw that but end of the day, the chips themselves don’t have hardware support for FP4. There’s smart ways around this limitation but it will never natively be close to true FP4 performance like MXFP4 and NVFP4 (happy to be proven wrong though).
- 3mo ago
- jmyeet 3mo agoWe're maybe only 2 years away from really useful, relatively affordable local LLM usage. You can buy a 5090 PC for $5-6k but 32GB of VRAM really limits model sizes to ~31B. And that won't change (even with NVidia's next generation) because NVidia uses VARM as an aggressive market segmentation technique. No, the hope really is these other platforms with a shared memory architecture. The DGX Spark won't be it because of the aforementioned market segmentation. So that leaves two players: AMD and Apple. The AMD platform is still too low memory bandwidth, currently <300GB/s. For comparison a 5090 (or 6000 Pro) is 1.8TB/s and the M3 Ultra Mac Studio is ~900GB/s. Oh and B100/B200 uses HBM3e memory at ~3.2TB/s. The M5 Max in some Macbook Pros tops out at ~600GB/s. So you need access to better RAM and better CPU architecture for all this. My great white hope is Apple. They have the market power to get memory and build silicon that coul dhave enough FLOPS to compete with NVidia's platform. They've started talking about it and I've seen rumors they're targeting the M7 generation (2028) for a huge leap. I'll believe it when I see it however. But the point is, I think we'll be running 31B models at 100+tok/s on enthusiast hardware in 2 years and we'll likely be able to locally run 100-400B models, possibly larger.
- jmward01 3mo agoIt seems like there is a very health space for an MOE targeted GPU where it has essentially an 5070ti with 16gb ish GDDR7 but then also has 128 GB LPDDR5x (or even just DDR5 as expansion dimms on it?). Putting this into the same card would likely reduce the transfer hit when a cache miss happened and the gpu had to load from the slower LPDDR5x. No need to have PCIe 5x16 limiting memory transfer if it is on the card. MOE models could then get near native performance and even models where the active parameters + context didn't quite fit the thrashing would be less of a problem. Not UMA but gets the UMA 'lots of system memory to play with' benefit.
- LoganDark 3mo agoSounds like Bolt's GPU
- jmward01 3mo agotheir main mem is LPDDR5x with DDR5 SODIMMs for expansion. LPDDR5x is 256GB/s? on the machines you see it implemented on which is a lot faster than the DDR5 expansion. For comparison, gddr7 on a 5070ti pumps ~900GB/s. 4x the tokens/sec if you are memory bound.
- khalic 3mo ago250GB/s on unified memory? That doesn’t sound right, it’s very low
- hedgehog 3mo agoWhat gets more?
- robotswantdata 3mo ago12 channel DDR5 6400, 614GB/s peak theoretical
- khalic 3mo agoFrom memory, a mbp m5 128gb is around 800
- speed_spread 3mo agoIt's still regular CPU memory, based on DDR. It's 256 bits wide but it's not GDDR. Unified just means that the CPU and GPU use the same physical memory bank. It does not make memory magically faster.
- owaislone 3mo agoIsn't this basically the same a Framework Desktop which has been out for quite a while? Does it improve on it in any way? https://frame.work/desktop https://frame.work/desktop
- wmedrano 3mo agoSeems like it https://www.phoronix.com/review/amd-ryzen-ai-halo/9 https://www.phoronix.com/review/amd-ryzen-ai-halo/9
- Grombobulous 3mo agoThe framework desktop also has a small advantage where you can buy it as a standard ITX board and purchasing it that way gives you access to a 4x PCIe slot. Buying it that way also kicks the price down a decent amount.
- t0mpr1c3 3mo agoMeh. The EVO-X2 came out six months and is cheaper. Any word yet on when the AMD Ryzen AI Max+ 495 is coming out?
- ronn00 3mo ago[flagged]
- temp2525 3mo agoamd ai apus are incapable of running models fast enough for productive work, you need over 100t/s for text mode and above 180t/s for image gen/ image recognition to not wait minutes. you cannot run claw and leave it - you would need to wait multile hours for small project like chatbot+landing page. price wise hw was not worth it even when it cost under 2k$, nowadays it cost even more. just a perspective(you can re check results in youtube reviews): ryzen ai runs most moe models under 60t/s while nvidia gpu can runt them at 100t/s. and you preferably need deep models for code, they would run at 40t/s.