13 ms·
Mistral ships Le Chat – enterprise AI assistant that can run on prem
- phupt26 1y agoAnother new model ( Medium 3) of Mistral is great too. Link: https://newscvg.com/r/yGbLTWqQ https://newscvg.com/r/yGbLTWqQ
- 85392_school 1y agoThis announcement accompanies the new and proprietary Mistral Medium 3, being discussed at https://news.ycombinator.com/item?id=43915995 https://news.ycombinator.com/item?id=43915995
- codingbot3000 1y agoI think this is a game changer, because data privacy is a legitimate concern for many enterprise users. Btw, you can also run Mistral locally within the Docker model runner on a Mac.
- kergonath 1y ago> I think this is a game changer, because data privacy is a legitimate concern for many enterprise users. Indeed. At work, we are experimenting with this. Using a cloud platform is a non-starter for data confidentiality reasons. On-premise is the way to go. Also, they’re not American, which helps. > Btw, you can also run Mistral locally within the Docker model runner on a Mac. True, but you can do that only with their open-weight models, right? They are very useful and work well, but their commercial models are bigger and hopefully better (I use some of their free models every day, but none of their commercial ones).
- distances 1y agoI also kind of don't understand how it seems everyone is using AI for coding. I haven't had a client yet which would have approved any external AI usage. So I basically use them as search engines on steroids, but code can't go directly in or out.
- trollbridge 1y agoMost my clients have the same requirement. Given the code bases I see my competition generating, I suspect other vendors are simply violating this rule.
- fhd2 1y agoYou might be able to get your clients to sign something to allow usage, but if you don't, as you say, it doesn't seem wise to vibe code for them. For two reasons: 1. A typical contract transfers the rights to the work. The ownership of AI generated code is legally a wee bit disputed. If you modify and refactor generated code heavily it's probably fine, but if you just accept AI generated code en masse, making your client think that you wrote it and it is therefore their copyright, that seems dangerous. 2. A typical contract or NDA also contains non disclosure, i.e. you can't share confidential information, e.g. code (including code you _just_ wrote, due to #1) with external parties or the general public willy nilly. Whether any terms of service assurances from OpenAI or Anthropic that your model inputs and outputs will probably not be used for training are legally sufficient, I have doubts. IANAL, and _perhaps_ I'm wrong about one or both of these, in one or more countries, but by and large I'd say the risk is not worth the benefit. I mostly use third party LLMs like I would StackOverflow: Don't post company code there verbatim, make an isolated example. And also don't paste from SO verbatim. I tried other ways of using LLMs for programming a few times in personal projects and can't say I worry about lower productivity with these limitations. YMMV. (All this also generally goes for employees with typical employment contracts: It's probably a contract violation.)
- distances 1y agoYes these are indeed the points. I don't really care too much, it would make me a bit more efficient but I'm billing by the hour anyway so I'm completely fine playing by the book.
- fhd2 1y agoNot sure I can agree with the "I'm billing by the hour" part. I mean sure, but I think of my little agency providing value, for a price. Clients have budgets, they have limited benefits from any software they build, and in order to be competitive against other agencies or their internal teams, overall, I feel we need to provide a good bang for buck. But since it's not all that much about typing in code, and since even that activity isn't all that sped up by LLMs, not if quality and stability matters, I would still agree that it's completely fine.
- ATechGuy 1y agoHave you tried using private inference that uses GPU confidential computing from Nvidia?
- Tepix 1y agopremises, not premise. https://www.grammar-monster.com/easily_confused/premise_premises.htm https://www.grammar-monster.com/easily_confused/premise_prem...
- demarq 1y agoAlso it’s like saying you can host a database on your Mac. Unless you have experience hosting and maintaining models at scale and with an enterprise feature set, then I believe what they are offering is beyond (for now) what you’d be able put up on your own.
- burnte 1y agoI have an M4 Mac Mini with 24GB of RAM. I loaded Studio.LM on it 2 days ago and had Mistral NeMo running in ten minutes. It's a great model, I need to figure out how to add my own writing to it, I want it to generate some starter letters for me. Impressive model.
- ulnarkressty 1y agoI think many in this thread are underestimating the desire of VPs and CTOs to just offload the risk somewhere else. Quite a lot of companies handling sensitive data are already using various services in the cloud and it hasn't been a problem before - even in Europe with its GDPR laws. Just sign an NDA or whatever with OpenAI/Google/etc. and if any data gets leaked they are on the hook.
- boringg 1y agoGood luck ever winning that one. How are you going to prove out a data leak with an AI model without deploying excessive amounts of legal spend? You might be talking about small tech companies that have no other options.
- nicce 1y ago> Btw, you can also run Mistral locally within the Docker model runner on a Mac. Efficiently? I thought macOS does not have API so that Docker could use GPU.
- jt_b 1y agoI haven't/wouldn't use it because I have a decent K8S ollama/open-webui setup, but docker announced this a month ago: https://www.docker.com/blog/introducing-docker-model-runner https://www.docker.com/blog/introducing-docker-model-runner
- nicce 1y agoHmm, I guess that is not actually running inside container/ there is no isolation. Some kind of new way that mixes llama.cpp , OCI format and docker CLI.
- v3ss0n 1y agoWhat's the point when we can run much powerful models now? Qwen3 , Deepseek
- _bin_ 1y agoIt would be short-termist for Americans or euros to use chinese-made models. Increasing their popularity has an indirect but significant cost in the long term. china "winning AI" should be an unacceptable outcome for America or europe by any means necessary.
- atwrk 1y agoWhy would that be? I can see why Americans wouldn't want to do that, but Europeans? In the current political climate, where the US openly claims their desire to annex European territory and so on? I'd rather see them prefer a locally hostable open source solution like DeepSeek.
- tigroferoce 1y agoMy two cents, as European, is that since we are more and more asking to LLMs for information, it wouldn't be wise to let a foreign country, not even truly democratic, to choose the information we get.
- jamesblonde 1y agoThe Chinese don't get any of information if we use self-hosted DeepSeek or Qwen. They are open-source. You can run them in an air-gapped environment that can't phone home.
- fennecbutt 1y agoBut their models are gimped by bad censoring. At least I can still ask chatgpt how many innocent civilians America has bombed.
- simonw 1y agoThere are plenty of other ways to run Mistral models on a Mac. I'm a big fan of Mistral Small 3.1. I've run that using both Ollama (easiest) and MLX. Here are the Ollama models: https://ollama.com/library/mistral-small3.1/tags https://ollama.com/library/mistral-small3.1/tags - the 15GB one works fine. For MLX https://huggingface.co/mlx-community/Mistral-Small-3.1-24B-Instruct-2503-8bit https://huggingface.co/mlx-community/Mistral-Small-3.1-24B-I... and https://huggingface.co/mlx-community/Mistral-Small-3.1-24B-Instruct-2503-4bit https://huggingface.co/mlx-community/Mistral-Small-3.1-24B-I... should work, I use the 8bit one like this: llm install llm-mlx llm mlx download-model mlx-community/Mistral-Small-3.1-Text-24B-Instruct-2503-8bit -a mistral-small-3.1 llm chat -m mistral-small-3.1 The Ollama one supports image inputs too: llm install llm-ollama ollama pull mistral-small3.1 llm -m mistral-small3.1 'describe this image' \ -a https://static.simonwillison.net/static/2025/Mpaboundrycdfw-1.png Output here: https://gist.github.com/simonw/89005e8aa2daef82c53c2c2c62207f6a#response https://gist.github.com/simonw/89005e8aa2daef82c53c2c2c62207...
- indigodaddy 1y agoSimon, can you recommend some small models that would be usable for coding on a standard M4 Mac Mini (only 16G ram) ?
- simonw 1y agoThat's pretty tough - the problem is that you need to have RAM left over to run actual applications! Qwen 3 8B on MLX runs in just 5GB of RAM and can write basic code but I don't know if it would be good enough for anything interesting: https://simonwillison.net/2025/May/2/qwen3-8b/ https://simonwillison.net/2025/May/2/qwen3-8b/ Honestly though with that little memory I'd stick to running against hosted LLMs - Claude 3.7 Sonnet, Gemini 2.5 Pro, o4-mini are all cheap enough that it's hard to spend much money with them for most coding workflows.
- codetrotter 1y agoHow about on an MacBook Pro M2 Max with 64GB RAM? Any recommendations for local models for coding on that? I tried to run some of the differently sized DeepSeek R1 locally when those had recently come out, but couldn’t manage at the time to run any of them. And I had to download a lot of data to try those. So if you know a specific size of DeepSeek R1 that will work on 64GB RAM on MacBook Pro M2 Max, or another great local LLM for coding on that, that would be super appreciated
- ATechGuy 1y agoWhy not use confidential computing based offerings like Azure's private inference for privacy concerns?
- lolinder 1y agoGame changer feels a bit strong. This is a new entry in a field that's already pretty crowded with open source tooling that's already available to anyone with the time and desire to wire it all up. It's likely that they execute this better than the community-run projects have so far and make it more approachable and Enterprise friendly, but just for reference I have most of the features that they've listed here already set up on my desktop at home with Ollama, Open WebUI, and a collection of small hand-rolled apps that plug into them. I can't run very big models on mine, obviously, but if I were an Enterprise I would. The key thing they'd need to nail to make this better than what's already out there is the integrations. If they can make it seamless to integrate with all the key third-party enterprise systems then they'll have something strong here, otherwise it's not obvious how much they're adding over Open WebUI, LibreChat, and the other self-hosted AI agent tooling that's already available.
- troyvit 1y ago> crowded with open source tooling that's already available to anyone with the time and desire to wire it all up. Those who don't have the time and desire to wire it all up probably make up a larger part of the market than those who do. It's a long-tail proposition, and that might be a problem. > I have most of the features that they've listed here already set up on my desktop at home I think your boss and your boss' boss are the audience they are going for. In my org there's concern over the democratization of locally run LLMs and the loss of data control that comes with it. Mistral's product would allow IT or Ops or whatever department to set guardrails for the organization. The selling point that it's turn-key means that a small organization doesn't have to invest a ton of time into all the tooling needed to run it and maintain it. Edit: I just re-read your comment and I do have to agree though. "game-changer" is a bit strong of a word.
- abujazar 1y agoActually you shouldn't be running LLMs in Docker on Mac because it doesn't have GPU support. So the larger models will be extremely slow if they'll even produce a single token.
- raxxorraxor 1y agoI think the the standard setup for vscode continue for ollama is already 99% of ai coding support I need. I think it is even better than commercial offerings like cursor, at least in the projects and languages I use and have tested it. We had a Mac Studio here nobody was using and it we now use it as a tiny AI station. If we like, we could even embed our codebases, but it wasn't necessary yet. Otherwise it should be easy to just buy a decent consumer PC with a stronger GPU, but performance isn't too bad even for autocomplete.
- thepill 1y agoWhich models are you using?
- dzhiurgis 1y agoHow many is many? Literally all of them use cloud services.
- Palmik 1y agoI really don't see the big deal. Gemini also allows on-prem in similar fashion: https://cloud.google.com/blog/products/ai-machine-learning/run-gemini-and-ai-on-prem-with-google-distributed-cloud https://cloud.google.com/blog/products/ai-machine-learning/r...
- victorbjorklund 1y agoWhy use this instead of an open source model?
- _mlbt 1y ago> our world-class AI engineering team offers support all the way through to value delivery.
- iamnotagenius 1y ago[dead]
- victorbjorklund 1y agoGuess that makes sense. But I'm sure they charge good money for it and then you could just use that money for someone helping you with an open source model.
- disgruntledphd2 1y agoPresumably one throat to choke logic applies here, particularly in Europe.
- deleted 1y ago[deleted]
- curiousgal 1y agoToo little too late, I work in a large European investment bank and we're already using Anthropic's Claude via Gitlab Duo.
- croes 1y agoIs there are replacement for the Safe Harbor replacement? Otherwise it could be illegal to transfer EU data to US companies
- _bin_ 1y agoThe law means don’t do what a slow moving regulator can and will prove in court. In this case, the law has no moral valence so I doubt anyone there would feel guilty breaking it. He may mean individuals are using ChatGPT unofficially even if prohibited nominally by management. Such is the case almost everywhere.
- jagermo 1y agoAI data residency is an issue for several of our customers, so I think there is still a big enough market for this.
- alwayseasy 1y agoYour bank sticks with any tech that comes out first? How is this a cogent argument?
- guerrilla 1y agoInteresting. Europe is really putting up a fight for once. I'm into it.
- resource_waste 1y agoExpected this comment. Mistral has been consistently last place, or at least last place among ChatGPT, Claude, Llama, and Gemini/Gemma. I know this because I had to use a permissive license for a side project and I was tortured by how miserably bad Mistral was, and how much better every other LLM was. Need the best? ChatGPT Need local stuff? Llama(maybe Gemma) Need to do barely legal things that break most company's TOS? Mistral... although deepseek probably beats it in 2025. For people outside Europe, we don't have patriotism for our LLMs, we just use the best. Mistral has barely any usecase.
- byefruit 1y agoYou are probably getting downvoted because you don't give any model generations or versions ('ChatGPT') which makes this not very credible.
- resource_waste 1y ago[flagged]
- dismalaf 1y agoIn your first comment you mentioned you used Mistral because of its permissive license (so guessing you used 7B, right?). Then you compare it to a bunch of cutting edge proprietary models. Have you tried Mistral's newest and proprietary models? Or even their newest open model?
- thrance 1y ago"patriotic Europeans" is an... interesting combination of words. I'd almost call it an oxymoron.
- _pdp_ 1y agoWhile I am rooting for Mistral, having access to a diverse set of models is the killer app IMHO. Sometimes you want to code. Sometimes you want to write. Not all models are made equal.
- binsquare 1y agoWell that sounds right up the alley of what I built here: www.labophase.com
- the_clarence 1y agoTbh I think the one general model approach is winning. People don't want to figure out which model is better at what unless its for a very specific task.
- sschueller 1y agoCouldn't you could place a very light weight model in front to figure out which model to use?
- the_clarence 1y agoMy guess is that this is basically what AI providers are slowly moving to. And this is what models seem to be doing underneath the surface as well now with Mixture of Experts (MoE).
- F-Lexx 1y agoGood idea. Then you could place another lighter-weight model in front of THAT, to figure out which model to use in order to find out which model to use. It,'s LLMs, all the way down.
- sReinwald 1y agoThat’s a perfectly valid idea in theory, but in practice you’ll run into a few painful trade-offs, especially in multi-user environments. Trust me, I'm currently doing exactly that in our fairly limited exploration of how we can leverage local LLMs at work (SME). Unless you have sufficient VRAM to keep all potential specialized models loaded simultaneously (which negates some of the "lightweight" benefit for the overall system), you'll be forced into model swapping. Constantly loading and unloading models to and from VRAM is a notoriously slow process. If you have concurrent users with diverse needs (e.g., a developer requiring code generation and a marketing team member needing creative text), the system would have to swap models in and out if they can't co-exist in VRAM. This drastically increases latency before the selected model even begins processing the actual request. The latency from model swapping directly translates to a poor user experience. Users, especially in an enterprise context, are unlikely to tolerate waiting for a minute or more just for the system to decide which model to use and then load it. This can quickly lead to dissatisfaction and abandonment. This external routing mechanism is, in essence, an attempt to implement a sort of Mixture-of-Experts (MoE) architecture manually and at a much coarser grain. True MoE models (like the recently released Qwen3-30B-A3B, for instance) are designed from the ground up to handle this routing internally, often with shared parameter components and highly optimized switching mechanisms that minimize latency and resource contention. To mitigate the latency from swapping, you'd be pressured to provision significantly more GPU resources (more cards, more VRAM) to keep a larger pool of specialized models active. This increases costs and complexity, potentially outweighing the benefits of specialization if a sufficiently capable generalist model (or a true MoE) could handle the workload with fewer resources. And a lot of those additional resources would likely sit idle for most of the time, too.
- iamnotagenius 1y agoMistral models though are not interesting as models. Context handling is weak, language is dry, coding mediocre; not sure why would anyone chose it over Chinese (Qwen, GLM, Deepseek) or American models (Gemma, Command A, Llama).
- amai 1y agoData privacy is a thing - in Europe.
- tensor 1y agoCommand A is Canadian. Also mistral models are indeed interesting. They have a pretty unique vision model for OCR. They have interesting edge models. They have interesting rare language models. And also another reason people might use a non-American model is that dependency on the US is a serious business risk these days. Not relevant if you are in the US but hugely relevant for the rest of us.
- tootie 1y agoI flip back and forth with Claude and Le Chat and find them comparable. Le Chat always feels very quick and concise. That's just vibes not benchmarks.
- crowcroft 1y agoI haven't used it much, but I did find Le Chat to be FAST in a way that I don't always get with ChatGPT.
- caseyy 1y agoThis will make for some very good memes. And other good things, but memes included.
- m-hodges 1y agoI love that "le chat" translates from French to English as "the cat".
- Jordan-117 1y agoAlso, "ChatGPT" sounds like chat, j’ai pété ("cat, I farted")
- layer8 1y agoMistral should highlight more in their marketing that it doesn’t make you fart.
- foobahhhhh 1y agoInstead it disobeys commands, uses up your resources then you find it never belonged to you in the first place.
- cryptonector 1y agoI came in to say this, and I was sure I'd be the first. This is so appropriate considering how ChatGPT -like all LLMs- hallucinates.
- debugnik 1y agoTheir M logo is a pixelated cat face as well.
- AceJohnny2 1y agoI wonder if they mean to reference the Belgian comic Le Chat by Philippe Geluck. https://en.wikipedia.org/wiki/Le_Chat https://en.wikipedia.org/wiki/Le_Chat
- I_am_tiberius 1y agoI really love using le chat. I feel much more save giving information to them than to openai.
- FuriouslyAdrift 1y agoGPT4All has been running locally for quite a while...
- deleted 1y ago[deleted]
- starik36 1y agoI don't see any mention of hardware requirements for on prem. What GPUs? How many? Disk space?
- tootie 1y agoI'm guessing it's flexible. Mistral makes small models capable of running on consumer hardware so they can probably scale up and down based on needs. And what is available from hosts.
- rowanajmarshall 1y agoI run a Mistral model on my phone!
- dr_kretyn 1y agoExplain more please? Is that a big phone/tiny laptop with long GPU connector? Is that a tiny model?
- Havoc 1y agoNot quite following. It seems to talk about features common associated with local servers but then ends with available on gcp Is this an API point? A model enterprises deploy locally? A piece of software plus a local model? There is so much corporate synergy speak there I can’t tell what they’re selling
- frabcus 1y agoThey mention Google Cloud Marketplace (not Google Cloud Platform), this seems to be their listing there: https://console.cloud.google.com/marketplace/product/mistralai/le-chat-enterprise?inv=1&invt=Abw1hw&project=pristine-valve-430713-n0&rapt=AEjHL4NQGl8N72-thyLEjwBYPMXctOLykCV2BMZBtynFzmgcIsyU1KuWXLegEhyneTXizPPFHtFRr02JxbS_H3nUjtayVJSTS-X_uIM4zGE9Y-CcHITLHWk https://console.cloud.google.com/marketplace/product/mistral... Which says: "Managed Services are fully hosted, managed and supported by the service providers. Although you register with the service provider to use the service, Google handles all billing." My assumption is that they're using Google Marketplace for discovery and billing, and they offer a hosted option or an on-prem option. But agreed, it isn't clear!
- tecleandor 1y agoLota of tools offer billing you via Google Marketplace or the AWS equivalent as: - it joins billing with other stuff - I guess it's easier to get approval - and more important (at least in our case), it allows you to reach your Google Cloud (or AWS) contract commitments of expense, and keep your discounts :)
- badmonster 1y agointeresting take. i wonder if other LLM competitors would do the same.
- amelius 1y agoI'm curious about the ways in which they could protect their IP in this setup.
- beernet 1y agoMistral really became what all the other over-hyped EU AI start-ups / collectives (Stability, Eleuther, Aleph Alpha, Nyonic, possibly Black Forest Labs, government-funded collaborations, ...) failed to achieve, although many of them existed way before Mistral. Congrats to them, great work.
- stogot 1y agoI’m wondering why. More funding, better talent, strategy, or something else?
- agumonkey 1y agoi'm an outsider but none of the startups mentioned above ever came to my ears. Mistral suddenly popped after openai/anthropic exploded, and they were rapidly described as the 3rd contender, with emphasis on technical merit. Maybe i was fooled though.
- danielbln 1y agoBlack Forest Labs are the makers of FLUX, which for a while was the best open image model available (and generally a pretty strong image model). That said, now with a wave of Chinese models and the advent of autoregressive image models, I'm not sure how much that will stay true.
- Palmik 1y agoIt feels to me they turned into a generic AI consulting & solutions company. That does not mean it's a bad business, especially since they might benefit from the "built in EU" spin (whether through government contracts, regulation, or otherwise). One can deploy similar solution (on-prem) using better and more cost efficient open-source models and infrastructure already. What Mistral offers here is managing that deployment for you, but there's nothing stopping other companies doing the same with fully open stack. And those will have the benefit of not wasting money on R&D.
- 1y ago
- deleted 1y ago[deleted]
- mxmilkiib 1y agothe site doesn't work with dark mode, the text is dark also
- qwertox 1y agoThis is so fast it took me by surprise. I'm used to wait for ages until the response is finished on Gemini and ChatGPT, but this is instantaneous.
- adamsiem 1y agoParsing email... The intro video highlights searching email alongside other tools. What email clients will this support? Are there related tools that will do this?