2 ms·
Should've called it Safestral. Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumab
by fastball 2mo ago
Should've called it Safestral.
Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.
- himata4113 2mo agoIt's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.
- maelito 2mo agoDo you have a reference explaining these costs ? Part by part.
- lucrbvi 2mo agoMistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training. The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to. [1]: https://poolside.ai/ https://poolside.ai/ [2]: https://poolside.ai/blog/introducing-laguna-s-2-1 https://poolside.ai/blog/introducing-laguna-s-2-1
- winterismute 2mo agoIsn't poolside a completely different company from Mistral?
- brendoelfrendo 2mo agoYes, the point being made is that poolside is able to train large models with limited resources, which means that Mistral should be able to compete in that space, as they have access to much greater resources than poolside. Mistral simply chooses not to.
- drob518 2mo agoAnd Poolside’s latest models (Laguna S 2.1) are pretty good (not frontier, but competitive with the tier 2 models). Which means that Mistral could certainly compete in that space.
- tjwebbnorfolk 2mo agodo you have a source for their GB300 count?
- Alpha3031 2mo agoThe 13,800 number seem like it would be quoted from their annoucement of the $830 million funding round they had in March. Here's CNBC's article on it: https://www.cnbc.com/2026/03/30/mistral-ai-paris-data-center-cluster-debt-financing.html https://www.cnbc.com/2026/03/30/mistral-ai-paris-data-center... Not sure if they would have received the full number yet, but it's been a few months so they certainly could have. Bit of a moot point when the comparison was against Poolside's Laguna which isn't really "general" SOTA but SOTA-for-the-size, and Mistral is clearly capable of training 700B or 120B models that are that when released considering they have done that... A 2-3T model is probably possible with the GPUs they have but they would need to spend most of their resources on it, and it's not clear why they would want to.
- himata4113 2mo agoFrom my experience their capabilities are extremely narrow and generally perform terribly when faced with issues outside comparatively narrow training data.
- moffkalast 2mo agoWhat, you don't want a model called the Shitstral-3B :D
- paradox460 2mo agoBet drummer could make one. He did that hilarious one that injected ads into copy
- moffkalast 2mo agoHaha yeah, rivermind. A drummer tune that flips it around entirely and only lets through unsafe stuff would be pretty hilarious ngl.
- Jumpstylish 2mo agoI think it’s encouraging that the major players in AI are all focusing on what they do best. The United States is focused on new, cutting edge technology. China is focused on improving and optimizing the process for maximum efficiency. Europe is focused on creating useless administrative overhead. Everyone is in their element.