5 ms·
OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasonin
by trouve_search 4mo ago
OK, I'm 100% rooting for both Mistral and task focused small models.
But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now.
Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter its size.
Back one year ago with Mistral Small 3.1 they were keeping up, but they've fallen into irrelevancy right now.
If Mistral seriously wants to play the on-prem and small task-specific model game, a decent proxy would be to build models that get the r/localLlama crowd excited
- echelon 4mo agoNobody trying to compete with Google, OpenAI, and Anthropic should be playing the small models / local models game. Foundation model labs should be building very large reasoning models, then leaving it to the community to distill them down. You can't scale a small model up, but you can scale a small model down. I'm convinced the only way we'll have a seat at the table in the future and avoid total runaway takeoff is if there are very large models within 80% of the capabilities of the frontier models. Tiny RTX models do diddly squat to remain competitive. Build open weights models for running on H200s. I'll spin them up on RunPod or Lambda.
- ahnick 4mo agoI thought distillation meant small models don't have to compete with the big models and can always eventually achieve close parity, but it's just a matter of time to do the distillation? (i.e. how much lag do you want to live with) Am I oversimplifying?
- gertlabs 4mo agoThere is likely a theoretical limit to how much intelligence you can pack into a model of a given size (especially when stretching that over a large input context size). Our evals are pretty complex so we only recently started testing ~30B class models, which are now becoming quite smart (on par with the frontier from 1 year ago). Mistral is far behind, but I'm rooting for them. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- farley13 4mo agoI do think there's a chance open weight models have a bit of a moment with the costs of frontier models growing on business balance sheets. It's unfortunate from my "privacy loving" PoV that it's mostly Chinese models filling the gap. ( the top models on openrouter for instance ). I have used Mistral models out of pure ideology for web agents and the like which aren't doing a lot of heavy lifting.
- theturtletalks 4mo agoAntirez’s Deepseek 4 Flash implementation that can run on MacBooks also was a revelation. It runs decently on M5 Max 128GB and it’s pointing out other bottlenecks like prefill speed which will improve.
- lettergram 4mo agoWe actually found the Mistral Small 4, quantized to 4bit was comparable to Qwen 3.6 27B and is roughly the same size. At least from our experience on our use cases, the quantization of the Mistral model worked far better than trying to quantize the Qwen family. Fully agree to your point though, Mistral in general is far behind where I'd expect and Qwen in particular is crushing it at the smaller sizes. Personally, I'd consider anything 20B params and above a "medium" model. Small being <20B and large >100B. I think obviously we can get to the huge 1-2T param models, but frankly the margin of accuracy improvement for the speed hit is kinda insane (1-2% for many metrics).
- rhdunn 4mo agoIt's all relative. For local use I'd classify it by hardware (VRAM size) using FP8 or Q6 quantization: 1. tiny <2-3B -- easily runnable on lower-spec hardware 2. small 4-8B -- runnable on 8GB GPUs 3. medium 9-12B -- runnable on 12GB GPUs 4. large 13-24B -- runnable on 16GB (for the lower end models) and 24GB GPUs 5. very large 25-32GB -- runnable on 32GB GPUs 6. huge >32GB -- not easily runnable on consumer GPUs without compromising performance (offloading layers to the CPU/RAM), quality (heavy quantization, esp. at <= Q4), or price (investing in multi-GPU setups and/or server-grade hardware). You could possibly split huge down further, as 70GB models (e.g. llama 3) are easier to get working than >120GB models and 1TB models are completely intractable.
- sroussey 4mo agoAs a Mac user: 1. tiny <2-3B -- could run in a browser even, mac neo 2. small 4-8B -- last of browser options, MacBook Air base 3. medium 9-24B -- 32GB machine, air or pro notebook or mini 4. large 25-48B -- 64GB, pro notebook or mini 5. x-large 49-100B -- 128GB MacBook Pro or Studio 6. Huge > 100B -- 256/512GB Mac Studio
- ElFitz 4mo ago> tiny <2-3B -- could run in a browser even, mac neo Or a phone. I’m running Gemma 4 E2B in one of my apps on my 14 pro (which may or may not be killing my display through overheating. It might just be a coincidence).
- baq 4mo agoagreed, the next price increase from frontier labs (and the inevitable limits decrease in subscription tiers) will have people thinking real hard about their model providers and that's when mistral should be ready. however, given their recent performance, I realistically don't have my hopes high up.
- djvdq 4mo agoAlso, new Medium 3.5 is far more expensive than previous Mistral models, and much more expensive than e.g. Deepseek
- KronisLV 4mo agoI tried it out on some dev tasks with their Mistral Vibe subscription, and the performance was pretty okay (okay, not great), both in regards to development and speed. Worse than Anthropic's models I'm used to but at 20 EUR per month it wasn't a bad deal - except that the 200k context size would more or less be a deal breaker in many cases.
- eisa01 4mo agoWhere do you sign up for that subscription? I wanted to try out Mistral, but I fail to find anything like that even after creating an account
- djvdq 4mo agoMaybe on their pricing page? https://mistral.ai/pricing/ https://mistral.ai/pricing/
- KronisLV 4mo agoThe other comment already mentioned that you get their subscription: https://mistral.ai/pricing/ https://mistral.ai/pricing/ they do say that you can try out their coding agent for free, but personally the Pro tier is pretty affordable too to try out for a month. Then you can install their coding harness, I personally used the Python + uv option: https://mistral.ai/products/vibe/code/ https://mistral.ai/products/vibe/code/ if you don't have uv yet, you might have to install it too: https://docs.astral.sh/uv/ https://docs.astral.sh/uv/ though I already use it for other projects. Oh and if on Windows, you probably want to do all of the installation inside of WSL, just so that file paths are the *nix variety, I've had issues otherwise with pretty much every coding harness, like OpenCode as well (across multiple models). After that, you need an API key for your subscription, you can generate and copy it here: https://console.mistral.ai/codestral/cli https://console.mistral.ai/codestral/cli that's also where you see the quota, though it seems to NOT refresh instantly, but more or less a few times a day. Either way, happy coding!
- greyskull 4mo ago> task focused small models This is tangential: and forgive my ignorance here, but is there an inherent reason why there aren't smaller, focused models from the frontier model providers? I'm thinking something like a software-specific subset of Opus that is the default for use in Claude Code. Smaller, cheaper to deploy and consume, maybe faster.
- pavpanchekha 4mo agoOpenAI used to make Codex-specific models, but they stopped. What I've gathered from interviews and similar is that training two models isn't worth the (small) lift from having a coding-specific model. You're pre-training on everything anyway, and coding RL is reasonably useful for general-purpose models too.
- greyskull 4mo agoInteresting. I'd have guessed there would be meaningful opex benefits to serving smaller models.
- deleted 4mo ago[deleted]
- mediaman 4mo agoWhat I've heard is that much of the model "intelligence" is a commingled bucket: although you can specialize specific knowledge somewhat, it's hard to specialize advanced reasoning to specific domains because so much of reasoning is a generalized capability that is not unique to, say, coding. It turns out coding has to do with a lot of the same reasoning needed in math or in legal analysis, even if the grammatical expression is different. This is less true of lower intelligence tasks. Classification requires a lot less reasoning capacity and so can be much smaller and more specialized.
- ar0 4mo agoI agree. I am a paying Le Chat Pro user, really rooting for a European alternative. But the quality difference between Mistral and the frontier labs is growing too big to ignore. It’s worrying to me that they didn’t talk much about new models at the conference, because that is really where their focus should be IMHO. I am wondering what is keeping them back, though: Money? Compute? Skills? Training data? My fear is that you are really only getting really good models by training on very dubious data (outputs from the frontier models etc) and that Mistral is too European and too enterprisey to take those risks.
- mattnewton 4mo agoMy theory with no insider information: it’s a little of all of the above, but mostly money. To some extent, you can dig yourself out of a data hole with RL and a lot of compute. And you can buy a lot of compute and some data with a lot of money. Big labs have been operating in this regime for a while and it’s one of the drivers behind their costs beyond just scaling the weights and doing the actual training. Mistral just doesn’t have access to this level of compute or the money to try and muscle their way in.
- MichaelZuo 4mo agoDon’t they supposedly have a huge amount of EU support? Or at least there’s been a lot of noise about that.
- baq 4mo agoThey can get what, 1B euros? 10B when everyone loses their mind? This doesn’t buy nearly enough compute nowadays. Meanwhile, Anthropic and OpenAI have investors practically begging them to let them buy this much equity at mind-bogging valuations.
- MichaelZuo 4mo ago[flagged]
- coredev_ 4mo agoI don't agree that they are falling behind. Using both chat and cli I get what I need and it's comparable to "sota" when I compare.
- rhdunn 4mo agoYeah. I run LLM models locally and for me 22B-32B is the largest I'm willing to invest in trying out. Even though Mistral 4 has 6B active parameters per token (allowing 3-3.5 per token parameters to be loaded on a 4090), the ~240GB download + storage is pushing the limits of being able to try this out locally, especially if you are downloading and evaluating multiple models. It also makes it harder for other people to make downstream finetunes like with what happened with the older Mistral/Magistral models.
- wolttam 4mo agoI think machines like the DGX Spark are about to become a lot more common/popular. It’s big enough to run sparse 150-250B MoEs with enough throughout for a single user. Deepseek v4 Flash is #1 (in terms of usage) on OpenRouter because it’s good enough to be useful. You can run it on a Spark (though it runs better across 2, which is getting up there in cost)
- dyauspitr 4mo agoMistral is bad bad. For its use cases I feel like India’s Sarvam is doing better.
- ctrlkctrls 4mo agochanneling Rocky (extraterrestrial) there I see :)
- kergonath 4mo ago> a decent proxy would be to build models that get the r/localLlama crowd excited I don’t really disagree with your post, but this is not exactly right. That subreddit seems to go from hype train to hype train every week, I haven’t found anything really insightful in it for quite a while now.
- chartpath 4mo agoI find Mistral Medium 3.5 with OpenCode is perfectly fine if you're willing to talk to it in a more fine-grained way about actual code. For me that's fine because even with huge frontier models I don't like trying to vibe prompt like a product manager.
- thatsadude 4mo agoNawh, they trained on test since Llama 2, no wonder.
- arkh 4mo agoMistral is entering the "let's extract has much money from EU taxpayers as we can" phase of European tech company which did not get bought by a US one. They'll end like Dailymotion, just a zombie company.
- raincole 4mo ago> they've fallen into irrelevancy right now It's a very charitable take, as Mistral has never really left the realm of irrelevancy. It's only a matter of time before EU falls back to hosting Chinese models in EU datacenters.
- barrell 4mo agoI think it really depends on what you’re doing. I use mistral for many tasks in https://phrasing.app https://phrasing.app and they blow models many times their size out of the water. None of my tasks use reasoning though (reasoning actually kills the performance) so perhaps that’s why. Still, I just had to rewrite my pipeline, and mistral was both faster, cheaper, and substantially better than any alternative