3 ms·
Time to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days). Any idea where this model sits according toquality be
by hedora 1mo ago
Time to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days).
Any idea where this model sits according toquality benchmarks? Pre-bubble MSRP on this hardware was $1400, and it draws 200-ish watts, putting it down into consumer territory.
I’m wondering if it can replace claude for llm-friendly coding tasks.
- cpburns2009 1mo agoSo back in the Qwen 3.5 release, the 122B-A10B model scored slightly better than the 27B model. I'd expect this new 125B-A6B to perform similarly to the recently released 27B. Qwen3.8 27B is supposed to rival Sonnet/Opus 4.6.
- hedora 1mo agoThanks. My current stack ranking of anthropic models is: 4.6 ~= 4.8 4.7 much worse. Fable and newer consistently tells me to pound sand, so I’m not sure what I’m paying $200/month for. 4.8 sometimes does too, but it’s at least usable most of the time. So, I’d expect this to mostly replace Claude for my workflows. The main tradeoff for me should mostly be token throughput vs. no longer really trusting anthropic.
- cyanydeez 1mo agoI've got the A10B hooked up to deer-flow and it does remarkable well when you dont need to baby sit it.
- hugmynutus 1mo agoQwen3.8/Qwen3.6 has a weird self doubt/thinking too much problem. You can prompt it away. I would say it "approximates" Opus 4.X class models well enough especially for coding/linux problems. The only reason I stopped using it as much is I was getting 25-35tok/s on Intel B70 (non-quant) which made some responses slow. For a long running/autonomous task, it would probably be sufficient.
- eightysixfour 1mo agoTwo things: - check your temp settings vs. Qwen's recommendations, they specify what it should be in the model card for thinking on, and that reduces some "over" think. - the model appears to be intentionally designed to do a lot more test-time compute, if you anthropomorphize the tokens, it looks like overthinking and anxiety, but it is just spending compute to get to the end result, so it may not actually help it to prompt down the token spend (depending on the problem)
- wongarsu 1mo agoThere is the rule of thumb that if you take the geometric mean of the total and active parameters of an MoE model you get the equivalent size of an equally capable dense model. If you follow that formula, you would expect a 125b-a6b model to match a 27b model (sqrt(125*6) = 27.3). That does not feel like a coincidence
- FuckButtons 1mo agoWhere does that rule of thumb come from?
- wongarsu 1mo agoThat's a great question. I learned it on HN. Some searching around suggests it originated as an empirical observation in the local LLM space around 2023-2024 It's obviously just a rough approximation. Actual scaling laws suggested in published papers are a lot more complex, and even then you run into issues (architecture changes, effects like better training, putting intelligence on a one-dimensional axis is stupid in the first place, etc). But as an approximation it holds up pretty well for normal-ish ratios between active and total parameters