8 ms·
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
- nezhar 2mo agoModel: https://huggingface.co/mistralai/Shieldstral-1.0-3B https://huggingface.co/mistralai/Shieldstral-1.0-3B
- lenerdenator 2mo agoI'd really like to see more conversation around Mistral's models. It's good to see Europe developing AI.
- BlackRabbit1 2mo agoThe problem is that their performance is too far away from the latest generation of Asian models. They had kept up in the mid-range a few years ago. But this standing is sadly long gone. If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.
- yborg 2mo agoBy this logic the Chinese should have just given up and let the American AI companies have the market because they were so far behind. I'm sure Europe has the capability to distill other people's frontier models to catch up if they wish to do so.
- baq 2mo agoDistilling is unsafe from export control perspective - Chinese models are poisoned by US frontier distillation and a case can be made that the US won’t like distilling what they may consider transitively theirs, which they will the moment you’re anywhere near competitive.
- LunaSea 2mo agoUS judges have already rules that output of an LLM can't be copyrighted so not sure what would prevent Chinese companies to use said output for distillation purposes.
- NitpickLawyer 2mo ago> already rules that output of an LLM can't be copyrighted Mind sharing such cases? I'm not aware of any so far. There's the one with images, but that's commonly miss-understood, that case was ruled on a technicality (i.e. copyright needs to be attributed to a person, not a model)
- Me1000 2mo agohttps://www.copyright.gov/newsnet/2025/1060.html https://www.copyright.gov/newsnet/2025/1060.html
- baq 2mo agoNote I didn't mention copyright
- kergonath 2mo agoBy which other mechanism could American AI companies prevent this? Other companies don’t really care about EULAs and even if they needed to care it’s trivial to let third parties do it. Why would they? Almost nobody in the space cares about copyright and play fast and loose with laws and regulations. What’s the mechanism that could today prevent other companies from using LLM outputs to train their models?
- grezql 2mo ago[dead]
- kergonath 2mo ago> Distilling is unsafe from export control perspective That is not the direction American judges are taking. Right now, they are saying that LLM output cannot be copyrighted. And if looting copyrighted works for training is fair game, I really don’t see how one could argue that learning from other LLMs is not.
- sroussey 2mo agoThe parent commenter was talking about Mistral as a single company and you switched from that to all of the EU. There definitely have been Chinese companies with models that fell behind, which is the more direct comparison. As for the EU in general, there are not a lot of known options. There are some working on things. The US, the EU, and China all have frontier labs that have yet to release anything.
- deleted 2mo ago[deleted]
- DarkNova6 2mo agoYou mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places. Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.
- maigret 2mo ago“Distilled”? I mean what model is not distilled from other data? The American models happily trained from my blog and social media data without any kind of rewards. If I can pay the inference I don’t see why I wouldn’t do this. Also, I have not seen proof that the US lab do not use other models for training either.
- DarkNova6 2mo agoI'm referring to the formal usage of the term "distillation", instead of the informal one that refers to training it on any data as a whole. https://www.geeksforgeeks.org/nlp/what-is-llm-distillation/ https://www.geeksforgeeks.org/nlp/what-is-llm-distillation/
- selectively 2mo ago[dead]
- cyanregiment 2mo agoI'm from nor cal but always liked Mistral. Mistral 7b is still one of the best free/open models you can run locally on a MacBook. So fast too.
- ranger_danger 2mo agoWhat newer models have you compared this against? Surely even quants of Qwen3.5 or higher would blow it out of the water. Also what is your definition of "best"?
- petcat 2mo agoMistral needs to abandon their Everything-stral branding. Getting kind of lame. "Shieldstral" is an awkward and bad name
- braiamp 2mo agoThat naming only works if you commit to the bit even when it doesn't make sense. That builds branding.
- simlevesque 2mo agoPeople complain when a product use a familiar name that might collide and there's also people complaining when they invent new words altogether. Naming things is hard.
- cyanregiment 2mo agoWas this one the last stral for you? The stral the broke the camel's back?
- whythismatters 2mo agoThe shortest stral has been pulled for you
- fooofw 2mo agoIt seems like you're just clutching at strals now
- vardalab 2mo agoSays who? I kind of like it.
- ranger_danger 2mo agoThere can be other valid perspectives than your own
- bee_rider 2mo agoTheir web chat UI seems to be called “Vibe.” I think? I’m not sure if that’s the name of the product or just what they decided to label it in the browser. I wonder if -strap is just what they call the actual models, which are meant to be run “under the hood” anyway, so not really part of the branding. But I wish they could commit to the bit fully and call everything -stral. It’s quirky and self aware to give your products silly names.
- fastball 2mo agoShould've called it Safestral. Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.
- himata4113 2mo agoIt's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.
- maelito 2mo agoDo you have a reference explaining these costs ? Part by part.
- lucrbvi 2mo agoMistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training. The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to. [1]: https://poolside.ai/ https://poolside.ai/ [2]: https://poolside.ai/blog/introducing-laguna-s-2-1 https://poolside.ai/blog/introducing-laguna-s-2-1
- winterismute 2mo agoIsn't poolside a completely different company from Mistral?
- brendoelfrendo 2mo agoYes, the point being made is that poolside is able to train large models with limited resources, which means that Mistral should be able to compete in that space, as they have access to much greater resources than poolside. Mistral simply chooses not to.
- mosura 2mo agoSomeone should use this to do the exact opposite of the intention: filter for “offensive” content, and boost it or collate it into a newsletter/email blast for people of culture. You have to give it to Mistral they do at least know what the market near them says they want right now. The great problem is in a few years of this that market won’t be worth anything. Edit to add, you could also add this to an AI workflow so as to produce content that walks right up to the line but doesn’t trigger it.
- petcat 2mo ago> they do at least know what the market near them says they want right now It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content. I guess they know that the EU AI Act, Chat Control, etc are going to cause a lot of companies to need this kind of compliance.
- lava_pidgeon 2mo agoSome of the best social media is heavily moderate. This includes HN and r/credible defense . With a Quiet transparent and cheap LLM I imagine a social media website where you can have good discussion about everything around the world it would be a game changer and on my to-do list.
- petcat 2mo ago> Some of the best social media is heavily moderate Heavily moderated by humans with discretion. Not AI chat bots following a rules engine.
- winwang 2mo agoWouldn't discretion "just" be a really good rules engine?
- colechristensen 2mo agoThe bulk of moderation work is things which are easy and obvious. The correct way to moderate is automation with certainty falling back to humans with discretion. The new frontier of moderation should be blocking illiterate comments, as in the commenter is replying as though they didn't read or read and didn't understand.
- hypfer 2mo agoI would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _truly_ as flexible as claimed? __ Maybe something like "Is this guy a corporate fraud that is going to waste my time with performative nonsense?" That would be the true test for a moderation model and I would be immensely impressed if it could manage to pull that off. ___ Edit: Looking at the paper though.. probably not. I suppose this is useful for B2B, which seems to be mistrals whole thing. Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin. Maybe opinions on those base datasets could occasionally differ more than the model can be steered.
- charcircuit 2mo agoIt sounds like it is. You have a set of moderation policies and then you evaluate the model 1 time per policy if it is violating it. Then you combine the results into a score you use for taking actions off of.
- nikcub 2mo ago> which seems to be mistrals whole thing They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche. Before the datacenter deals their revenue was higher than xAI's There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost. Mistral, Microsoft model releases and Thinking Machines are all over this, and it's smart. Scoop up all the tasks that don't require large and expensive frontier general-purpose llms.
- bunyak 2mo ago[flagged]
- amos-burton 2mo ago[dead]
- elianaive 2mo agoI'm a bit doubtful that a black box approach like this to moderation will ever catch on.
- MagicMoonlight 2mo ago[dead]
- sroussey 2mo agoIt need only be one of several tools in a toolbox.
- pwython 2mo agoI've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.
- kergonath 2mo agoI am not sure how reliable it is in the real world. Also, in terms of liability, I don’t know how effective it would be to satisfy various regulations compared to a human moderator team.
- pwython 2mo agoI hear ya, but one could set different operating thresholds: auto-approve low-risk posts, hold ambiguous posts for review, and automatically reject very high-confidence violations. So HITL for sure, but MUCH less H in the L.
- FridgeSeal 2mo agoHuman-somewhere-nearish-the-loop!
- deleted 2mo ago[deleted]
- sbinnee 2mo agoYes it does look like a good solution. But when I imagine actually using a guardrail for a product, this model only outputs yes/no probabilities. There is no reasoning trace why it was rejected. Users or even developers would have no idea why a prompt was classified yes or no. I really like this release but I feel like I need something more to use it as a guardrail in production.
- xp84 2mo agoI think IRL in the “rejection” case they don’t want to tell the user exactly why, since the user may be malicious and use it to try to evade the block. And for use in moderating UGC, well, most platforms don’t take seriously the idea that they need to answer to their users. Only their advertisers. In the case of wondering why a bad thing got through, well, I think that’s why they just set these to the most pro-censorship level they can, to make that highly unlikely.
- snovv_crash 2mo agoFinally an AI company besides DeepSeek taking economics into account.
- storus 2mo ago[flagged]
- trilogic 2mo agoThis model is way small for a proper assessment (imo). It should be very useful to study how big the real model must be for this purpose. Maybe merging it to a bigger one (adding it as expert style in moe) would be a solution! Great job to Mistral team.
- gizmodo59 2mo agohow does this compare with https://developers.openai.com/api/docs/models/omni-moderation-latest https://developers.openai.com/api/docs/models/omni-moderatio... As for use cases, obviously we can't fully rely on non-deterministic capability for sensitive things but a small model which can do a good job acts as a first defense and then a human can review later.
- db29a0dbcd3b 2mo ago[flagged]
- gogasca 2mo ago[dead]
- theplumber 2mo agoSo a censorship model
- w4yai 2mo agoThanks Mistral !
- LAC-Tech 2mo agoOh so THAT'S why the web app is so slow
- TacticalCoder 2mo ago> ... a single yes/no question, e.g. "Does this content promote physical violence?" Is it honest about religious texts? Can I throw at it religious texts and it'll honestly tell me whether the text promotes physical violence or not?
- baby 2mo agoIt’s not American so maybe
- nc55g3g 2mo agoTried the demo. Works okay for basic stuff. Prompt-based policy is clever but I'm skeptical about real-world edge cases.
- deleted 2mo ago[deleted]
- CurbStomper 2mo ago[dead]
- ankushdograuk 2mo agothe strategy of focusing on smaller fine tuned models, from mistral is intresting
- ygouzerh 2mo agoCrazy that it's a small lab becoming the frontier in term of moderation models, instead of Meta which is pouring dozens of billions into LLMs. Meta would really benefit from work done on this front, however their model Llama Guards are quite lagging compared to the competition.
- porridgeraisin 2mo agoPretty funny that this is the thing Europe model is sota in.
- arttaboi 2mo ago[dead]
- hchja 2mo agoThe fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.
- xp84 2mo agoThat’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened. 2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system. 2026 scenario: all contact information has been scrubbed from the site. Users can click “chat” and a chatbot will apologize for their dissatisfaction and offer no option to escalate. No one will ever hear anything about the complaint, so there’s no need to explain the failure. User can either accept this or can get f**ked because all competitors operate the same way.
- nextaccountic 2mo agoNot all platforms want to operate like mainstream social media though
- maigret 2mo agoIn Germany you are obligated to provide usable contact information, and there are even lawyers who make their business model on suing you for not applying that perfectly (it’s has been abused a lot in the past decades btw).
- inigyou 2mo agoWhat about in France?
- tomjen3 2mo agoThis model is European. There are quite a lot of instances where under GDPR, you have the right to have incorrect information about you corrected. I also think you have the right to appeal a decision to a human (I got that message from Reddit once, because the bot could not understand the difference between discussion of the death penalty and threats to humans). There are also new rules about AI and what it can be used for. Mostly this restricts the government from AI-enhanced surveillance, which is good. But there are also issues regarding job security and automatically categorizing people based on AI. So this is super useful, but has potential issues depending on how it is deployed.
- 1saadcodes 2mo agoI'm liking the trend of companies are releasing smaller, focused models instead of trying to make one model do everything. A dedicated moderation model is much easier to reason about than hideden safety logic inside a general-purpose model which might not have had much training in that aspect at all
- blackqueeriroh 2mo agoFolks should check out https://zentropi.ai https://zentropi.ai and their latest model, CoPE-B-A4B: https://huggingface.co/zentropi-ai/cope-b-a4b https://huggingface.co/zentropi-ai/cope-b-a4b Policy adaptive models really are the coolest things these days. Also, check out https://roost.tools https://roost.tools for even more open safety tooling!
- 60pfennig 2mo agoThis is propably the sustainable future of AI. Small, efficient Models for narrow tasks.
- sudeepsd__ 2mo ago[dead]
- peri-cl 2mo agoI fed this model (Q8) the first chapter of Voltaire's Treatise on Tolerance and it says that it promotes violence against protected groups, <Instruct>: Given a query about the content, determine if the message meets it <Query>: Does this content promote violence against a protected group? <Document>: TRAITÉ SUR LA TOLÉRANCE, À l’occaſion de la mort de Jean Calas. CHAPITRE PREMIER. Hiſtoire abrégée de la mort de Jean Calas. LE meurtre de Calas, commis dans Toulouſe avec le glaive de la Juſtice, le 9me Mars 1762, eſt un des plus ſinguliers événements qui méritent l’attention de notre âge & de la poſtérité. On ... (truncated) yes
- xhkkffbf 2mo agoDoes the model only care about violence against "protected groups"? What about the people who aren't in those groups?
- alex_duf 2mo agocould the long s `ſ` be throwing the model off?
- peri-cl 2mo agoInteresting hypothesis. Replacing "ſ" with "s" did not change the output. I think the simple explanation is the likely one (the reason I deliberately chose this specific benchmark): the model isn't intelligent enough to figure out use/mention distinctions. It understands Voltaire is discussing injustice, violence, tolerance; but it doesn't understand which side he's on.
- alex_duf 2mo agothanks for testing it! I guess that would make sense for such a small model to be missing the subtlety
- Aachen 2mo agoTo save anyone else looking for it: the post doesn't mention multilingualism but Huggingface (https://huggingface.co/mistralai/Shieldstral-1.0-3B https://huggingface.co/mistralai/Shieldstral-1.0-3B) has a menu at the top where it specifies that it should understand French
- vvpan 2mo agoNot hotdog.
- spate141 2mo agoMade some notes and a working notebook for myself. If anyone wants to check out: - https://snehal.ai/shieldstral-policy-adaptive-moderation/ https://snehal.ai/shieldstral-policy-adaptive-moderation/ - https://github.com/spate141/latent-lab/tree/main/shieldstral https://github.com/spate141/latent-lab/tree/main/shieldstral
- YUART 2mo ago[dead]
- chris_explicare 2mo ago[dead]