2 ms·
Vulkan or ROCm backend? I've also got a strix halo box. 30tok/s would we usable, but I wonder how it compares to the Qwen3.8-Flash-next - I get about 40tok/s r
by UncleOxidant 8d ago
Vulkan or ROCm backend?
I've also got a strix halo box. 30tok/s would we usable, but I wonder how it compares to the Qwen3.8-Flash-next - I get about 40tok/s running that on Halogen and it feels like using Claude 4.6.
- syntaxing 8d agoVulkan, I have never used ROCm on it but have been debating since the latest big update. How is your prefill? Do you hit over 1K? If it’s 1000K prefill, and 40 TG, I might have to try this over the weekend. Also, can you fit 128K without offload the ngram onto SSD?
- UncleOxidant 8d agoI have not measured pre-fill, but it's said to be around 1000. It feels very snappy and unlike my experience with running 27B models the performance stays pretty flat even as the context increases. Unfortunately, we don't know how Halogen is doing this because it's closed source, but I think AMD should offer that guy some $$$ because he's done a lot of good work getting more performance out of Strix Halo.
- syntaxing 8d agoHalogen as in this right? https://github.com/peonist-ai/halogen-flash-server https://github.com/peonist-ai/halogen-flash-server
- UncleOxidant 8d agoyes
- syntaxing 8d agoThanks! I really like how the author packaged everything into a container. Definitely going to give it a go over the weekend!