8 ms·
I'm not sure what people are on in the comments. It doesn't beat the other models, but it sure competes despite its size. GLM 5.1 is an excellent model, but ev
by simjnd 5mo ago
I'm not sure what people are on in the comments. It doesn't beat the other models, but it sure competes despite its size.
GLM 5.1 is an excellent model, but even at Q4 you're looking at ~400GB.
Kimi K2.5 is really good too, and at Q4 quantization you're looking at almost ~600GB.
This model? You can run it at Q4 with 70GB of VRAM. This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM for ~3500 USD).
For the Claude-pilled people, I don't know if you only run Opus but when I was on the Pro plan Sonnet was already extremely capable. This beats the latest Sonnet while running locally, without anyone charging you extra for having HERMES.md in your repo, or locking you out of your account on a whim.
Mistral has never been competitive at the frontier, but maybe that is not what we need from them. Having Pareto models that get you 80% of the frontier at 20% of the cost/size sounds really good to me.
- redrove 5mo agoIt’s 128b dense model. Good luck getting more than 3t/s out of a mac. It doesn’t matter if it fits or not.
- zozbot234 5mo agoYou could run it on a single Mac Studio with M3 Ultra, or two Mac Studios with M4 Max at higher perf than that. And lightly quantizing this could give us modern dense models in the ~80GB size range, which is a very compelling target.
- freakynit 5mo agoWouldn't matter much still. M3 ultra has 819GB/s unified memory bandwidth. That means theoretical max tokem rate is 819/128 =~ 6.39 t/s. At 80 GB (5 bit quantization), its still near about 10 t/s ... far from a good coding experience. Also, these are theoretical max.. real world token generation rates would be at least 15-20% less.
- gregsadetsky 5mo agoI didn't know about HERMES.md ... (??) - found information here for others who are curious https://github.com/anthropics/claude-code/issues/53262 https://github.com/anthropics/claude-code/issues/53262
- giancarlostoro 5mo agoThat is insane, if you billed me an extra $200 for a bug in your system I'd flat out cancel my subscription. If you're not going to credit that back to me, you don't deserve anymore of my money. I'm a Claude first guy, but if you're going to bill me incorrectly, that's on you, own it, fix it.
- xcrjm 5mo agoThey did credit it back to him. There's a comment in the linked issue.
- MarsIronPI 5mo agoWhere? Just searched the entire thread for both the word "refund" and the word "credit" and I'm seeing nothing about credit being issued. Also what's with @sasha-id talking to himself? Looks weird as all get out.
- deleted 5mo ago[deleted]
- simjnd 5mo ago
- YetAnotherNick 5mo agoIt has similar SWE bench score to qwen 3.6 27b[1]. No one is comparing it to frontier. [1]: There is no other common benchmark in the blog.
- simjnd 5mo agoThat's more a testament of how good Qwen3.6 27B is (it really is great) more than how bad this one is IMO. Gemma 4 31B was already good, but Qwen3.6 27B is incredible for its size.
- reissbaker 5mo agoGood models vs bad models are relative: if this was released in 2020 it would be earth shattering. But releasing a model today that's only on par with open-source dense models a quarter of the size and soundly beaten by open-source MoEs with active param counts a quarter of the size is kind of a flop. The niche for this is basically no one. It'll run at near-zero TPS for the few local model aficionados with enough hardware to try it out, and is lower throughout and lower quality for people trying to use it at scale. I'm rooting for Mistral, I want them to release good models. This just isn't one. It's a little sad since they once were so prominent for open-source. Who knows — if they have the compute to train this, they have the compute to train an MoE that's 3-4T total params with 128B active. Maybe they'll make a comeback (although using Llama 2 attention is... not promising). I hope they do.
- freakynit 5mo agoI was hoping a lot from it... but this one, is not up to that mark. For example, here is it's comparion with 4.7x smaller model, qwen3.7-27b. https://chatgpt.com/share/69f239e8-7414-83a8-8fdd-6308906e5f73 https://chatgpt.com/share/69f239e8-7414-83a8-8fdd-6308906e5f... Tldr: qwen3.6-27b, a 4.7x smaller model, have similar performance.
- r0b05 5mo agoThat's a chatgpt summary. Actual usage would a better test.
- freakynit 5mo agoyep.. until then, this is good enough since the tests are standard, and the results are numeric and can be compared without any doubt.
- lostmsu 5mo agoTo be fair MoE from Qwen itself had the same "problem". 3.5 122B MoE was same or worse than 3.5 27B. Yet to see 122B 3.6. UPD. NVM, Mistral Medium 3.5 is dense. So yes, it is worse in every way.
- Aurornis 5mo ago> This model? You can run it at Q4 with 70GB of VRAM. This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM for ~3500 USD). The one thing I would want everyone curious about local LLMs to know is that being able to run a model and being able to run a model fast are two very different thresholds. You can get these models to run on a 128GB Mac, but we need to first tell if Q4 retains enough quality (models have different sensitivities to quantization) and how fast it runs. For running async work and background tasks the prompt processing and token generation speeds matter less, but a lot of Mac Studio buyers have discovered the hard way that it's not going to be as responsive as working with a model hosted in the cloud on proper hardware. For most people without hard requirements for on-site processing, the best use case for this model would be going through one of the OpenRouter hosted providers for it and paying by token. > This beats the latest Sonnet while running locally Almost every open weight model launch this year has come with claims that it matches or exceeds Sonnet. I've been trying a lot of them and I have yet to see it in practice, even when the benchmarks show a clear lead.
- zozbot234 5mo agoCloud hardware is not inherently more "proper" than what's being proposed here, there's nothing wrong per se about targeting slower inference speeds in an on prem single-user context.
- Aurornis 5mo ago> Cloud hardware is not inherently more "proper" than what's being proposed here Cloud hardware can run the original model. Quantization will reduce quality. The quality drop to Q4 is not trivial. Cloud hardware is also massively faster in time to first token and token generation speed. > there's nothing wrong per se about targeting slower inference speeds in a local single-user context. If that's what the user wants and expects then it's fine Most people working interactively with an LLM would suffer from slower turns.
- zozbot234 5mo ago> Cloud hardware can run the original model. Quantization will reduce quality. New models are often being released in quantized format to begin with. This is true of both Kimi and the new DeepSeek V4 series. There is no "original model", the model is generated using Quantization Aware Training (QAT).
- 2ndorderthought 5mo agoThe point is it's open weight and is tiny compared to a lot of it's competitors. 4gpus for world class performance - sweet!
- DeathArrow 5mo ago>This model? You can run it at Q4 with 70GB of VRAM. >This beats the latest Sonnet while running locally Not sure it will beat Sonet at Q4. >This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM for ~3500 USD). For $3500 I can get 7-8 years of GLM using coding plans, have a faster model and much better code quality.
- kobalsky 5mo ago> For $3500 I can get 7-8 years of GLM mind sharing where's the go to place to pay for open models?
- DeathArrow 5mo agoYou can get GLM coding plans from Z.ai and Ollama Cloud and OpenCode Go.
- simjnd 5mo agoI recommend using OpenRouter (openrouter.ai). Basically a broker between inference providers and you which allows you to pick, try, and switch models from a massive catalog, extremely transparent about usage and pricing.
- pbgcp2026 5mo ago+5% to every API call.
- rsanek 5mo agoI've had a decent experience with ollama cloud. It is slower than going thru openrouter but much, much cheaper -- the generosity of their $20 plan reminds me of what the Claude Code $20 plan was back in the day
- simjnd 5mo ago> Not sure it will beat Sonet at Q4. Very valid. Importance-weighted quantization and TurboQuant on model weights can reduce loss a lot compared to "traditional" Q4 so one can be hopeful. > For $3500 I can get 7-8 years of GLM using coding plans, have a faster model and much better code quality But you will own no computer, and that's also assuming prices stay what they are. Anyway my point was not whether or not it makes financial sense for everyone. A lot of people are very happy not owning their movies, software, games, cars or house. I'm just happy there is a future where the people can own and locally run the tech that was trained on their stolen data.
- liuliu 5mo agoThe competition is on DeepSeek v4 Flash for similar size / deployment target.
- simjnd 5mo agoDeepSeek v4 Flash is still over 100GB at Q4 IIRC, and Q4 has generally been the sweet spot. Although it's an MoE so it might run a lot faster that this dense Mistral model if you have the RAM.
- pbgcp2026 5mo ago"Q4 has generally been the sweet spot" for self-hosting, yes. For any real meaningful work it's dumb AF. The only way to get reasonable intelligence from mid-size Gemma or Qwen is to run full precision BF16. Anything else is just an emulation of AI.
- EntityDeletr 5mo agoI would disagree. I have 8 GB of VRAM and 32 GB of RAM. I can either run a 4B BF16 dense model fully on GPU at around 30 t/s or Qwen3.6 35B A3B Q5_K_M at 20 t/s with GPU offload. Which one would I choose?
- giancarlostoro 5mo ago> For the Claude-pilled people, I don't know if you only run Opus but when I was on the Pro plan Sonnet was already extremely capable. Before February I was able to use Opus on High exclusively on my Max plan no problem. Now I've shifted to just using Sonnet on high and yeah, its pretty capable. I love that, Claude Pilled. ;)
- simjnd 5mo agoYeah I love Claude, amazing models. Anthropic has very quickly burned most of the goodwill I had for it so I still ended up cancelling my subscription.
- UncleOxidant 5mo agoYeah, you can run it locally if you have enough VRAM, but the reports trickling in are saying about 3 tok/sec. This was on a Strix Halo box which definitely has the needed VRAM, but isn't going to have as high mem bandwidth as a GPU card, it's going to be similar on a Mac - that's the dilemma... the unified memory machines have the VRAM, but the bandwidth isn't great for running dense models. This size of a dense model is only going to be runnable (usefully) by very few people who have multiple GPU cards with enough memory to add up to about 70GB.
- simjnd 5mo agoI don't think this is quite correct, a Strix Halo box usually has 256 GB/s memory bandwidth. An M5 Max has 614 GB/s. An M3 Ultra (no M4 or M5 Ultra) has 820 GB/s. It's still not GDDR or HBM territory, but still significantly faster. That's the edge of Apple Silicon for AI. When they scale up the chip they add more memory controllers which adds more channels and more bandwidth. But yeah in the end it's still going to be only a handful of people that can run it. What I meant is that I think researching and developing smaller more powerful model is more interesting than chasing the next 3T parameter model while burning through VC money and squeezing your customer base more and more aggressively.
- zackangelo 5mo agoIsn't Kimi K2.6 natively INT4?
- simjnd 5mo agoI don't think any models are natively INT4? I wouldn't see the point to nerf the model out-of-the-box.
- zozbot234 5mo agoIt's not nerfed, it's natively trained at that quantization a.k.a. Quantization Aware Training.
- pbgcp2026 5mo agoQAT typically uses BF16/FP32 during the training process to simulate lower precision.
- EntityDeletr 5mo agoThe only model I have seen like that is GPT OSS, natively quantized to MXFP4.
- deepsquirrelnet 5mo agoI would love to be able to run frontier locally, but I think the larger importance of open weight models is price accountability. In the US with our broken system of capitalism, it’s the only way we can tether these companies to reality. Left to their own devices, I’m not convinced they would actually compete with each other on price. Buy nobody like to talk about how “moat” building is fundamentally anti-competitive, even in name. Funny that self proclaimed capitalists hate the system in practice. Commodity pricing is what truly terrifies them.
- simjnd 5mo agoI'm not necessarily interested in having frontier locally. You don't need to be frontier to be a very good and useful coding agent. I agree with your point on price accountability though. Hopefully no tariff comes down on the Chinese and European open-weight models.
- sayYayToLife 5mo ago[dead]
- ksubedi 5mo agoLet's not forget Qwen 35B A3B MoE. It gets better performance than this in all the metrics for a fraction of the memory / compute footprint. Sad to see all the non Chinese open source models being at least one generation behind.
- simjnd 5mo agoQwen3.6 27B is even more impressive IMO. Dense so it doesn't run as fast but it's so good.
- trueno 5mo agoim kinda torn on which to download. i have the headroom to run either, mostly just want the occasional "do a coding thing im too lazy to do"
- EntityDeletr 5mo agoThen go with Qwen3.6 35B A3B. It's way faster (up to 5x) and it is 80% as capable as the 27B. The 27B is for serious people looking for one shot coding. The 35B is for iterative and quicker coding. I am in the same situation as you (making something I don't want to do myself) and I use the 35B at Q5_K_M.
- revolvingthrow 5mo agoEh. Those results would be noteworthy if it was a a MoE. A 120B dense? Firmly in meh territory.
- gregorygoc 5mo agoWhy do you care?
- WhitneyLand 5mo ago“This beats the latest Sonnet while running locally” Not really. - The benchmarks are based on F8_E4M3 and you’re not running that on any Mac. - Sonnet has a 1M token context window. This is 256k but again you’re probably not even getting that locally. - Sonnet is fast over the wire. This is going to be much slower.
- trueno 5mo agothe benchmarks we're using to measure llm's do no justice when everyone's mental-benchmark is simply "is it going to feel like using claude" and the answer is still no. the entire llm space is stuffed with tons of crazy datapoints and vernacular that barely paint the picture of the mental benchmark everyone is after. i too am desperate to just sever ties with these big providers, my fingers are crossed we get there within the constraints of local hardware even if that means me spending 3-5k i just want off this wild ride.
- trvz 5mo ago> Sonnet is fast over the wire. Except when it’s unavailable. For sovereignity, the downsides are worth it to some.
- varispeed 5mo agoNot sure if 1M token window is meaningful with Sonnet/Opus. The models go dumb quickly as context increases making them unusable (that is if you get routed to actual Opus, otherwise they are just dumb regardless of context window).
- varispeed 5mo ago> It doesn't beat the other models, but it sure competes despite its size. But what is the rationale for running a dumb model? Because it can ocasionally produce something passable? I don't get where is the value apart from mild entertainment, as in "I am somewhat of Anthropic myself".
- simjnd 5mo agoAre you dumb because you're not Einstein? Intelligence is a spectrum. Just because you're not #1 doesn't mean you're dumb. A lot of small models are not frontier but are still very competent and are very useful coding agent. It may take better prompting and more guiding, but that can be a reasonable tradeoff for some people.