11 ms·
Leanstral: Open-source agent for trustworthy coding and formal proof engineering
Lean 4 paper (2021): https://dl.acm.org/doi/10.1007/978-3-030-79876-5_37 https://dl.acm.org/doi/10.1007/978-3-030-79876-5_37
- leontloveless 7mo ago[dead]
- sunwukong666 7mo ago[dead]
- blurbleblurble 7mo agoTruly exciting
- andai 7mo agoTrustworthy vibe coding. Much better than the other kind! Not sure I really understand the comparisons though. They emphasize the cost savings relative to Haiku, but Haiku kinda sucks at this task, and Leanstral is worse? If you're optimizing for correctness, why would "yeah it sucks but it's 10 times cheaper" be relevant? Or am I misunderstanding something? On the promising side, Opus doesn't look great at this benchmark either — maybe we can get better than Opus results by scaling this up. I guess that's the takeaway here.
- DrewADesign 7mo agoIt’s really not hard — just explicitly ask for trustworthy outputs only in your prompt, and Bob’s your uncle.
- miacycle 7mo agoAssuming that what you're dealing with is assertable. I guess what I mean to say is that in some situations is difficult to articulate what is correct and what isn't depending in some situations is difficult to articulate what is correct and what isn't depending upon the situation in which the software executes.
- DrewADesign 7mo agoAnd Bob’s your uncle.
- flowerbreeze 7mo agoThey haven't made the chart very clear, but it seems it has configurable passes and at 2 passes it's better than Haiku and Sonnet and at 16 passes starts closing in on Opus although it's not quite there, while consistently being less expensive than Sonnet.
- andai 7mo agoOh my bad. I'm not sure how that works in practice. Do you just keep running it until the tests pass? I guess with formal verification you can run it as many times as you need, right?
- ainch 7mo agopass@k means that you run the model k times and give it a pass if any of the answers is correct. I guess Lean is one of the few use cases where pass@k actually makes sense, since you can automatically validate correctness.
- teekert 7mo agoI also don't understand the focus on vibe coding in the marketing. Vibe coding kind of has the image of being for non-devs, right? I do like agents (like Claude Code), but I don't consider myself to be vibe coding when I use them. Either I'm using a language/framework I know and check every step. OR I'm learning, checking every step and asking for explanations. I tried vibe coding, and really dislike the feeling I have when doing it. It feels like building a house, but without caring about it, and just using whatever tech. Sure I may have moisture problems later, but it's a throwaway house anyway. That's how I feel about it. Maybe I have a wrong definition. Maybe it's good to not use "vibe coding" as a synonym for programming with agent assistance. Just to protect our profession. Like: "Ah you're vibing" (because you have Claude Code open), "No, I'm using CC to essentially type faster and prevent syntax errors and get better test coverage, maybe to get some smart solutions without deep research. But I understand and vouch for every loc here. 'We are not the same.'"
- DANmode 7mo ago> It feels like building a house, but without caring about it, and just using whatever tech. So, most homebuilders (in the US) unfortunately.
- teekert 7mo agoI myself am now and expert at insulation and all the vapor-permeable and vapor-blocking membranes/foils/foams that come with it. It came at great cost though, I hated the process of learning and the execution. I was less than happy for some years. But I feel even more uncomfortable vibe-home-improving than I do vibe-coding. The place is starting to look nice now though.
- benterix 7mo ago> I tried vibe coding, and really dislike the feeling I have when doing it. It feels like building a house, but without caring about it, and just using whatever tech. Sure I may have moisture problems later, but it's a throwaway house anyway. That's how I feel about it. Maybe I have a wrong definition. No, I feel the same. I vibe-coded a few projects and after a few weeks I just threw them away, ultimately I felt I just wasted my time and wished I coudl get it back to do something useful.
- lefrenchy 7mo agoDoes Mistral come close to Opus 4.6 with any of their models?
- DarkNova6 7mo agoNot at the moment, but a release of Mistral 4 seems close which likely bridges the gap.
- re-thc 7mo agoMistral Small 4 is already announced.
- androiddrew 7mo agoMOE but 120B range. Man I wish it was an 80B. I have 2 GPUs with 62Gib of usable VRAM. A 4bit 80B gives me some context window, but 120B puts me into system RAM
- Aerroon 7mo agoEither some q3 or since it's a MoE, maybe a REAP version of q4 might work (or could be terrible, I'm not sure about REAP'd models).
- tjwebbnorfolk 7mo agoMistral hasn't been in the running for SOTA for quite awhile now
- chucky_z 7mo agoI use mistral-medium-3.1 for a lot of random daily tasks, along with the vibe cli. I'd state from my personal opinion that mistral is my preferred 'model vendor' by far at this point. They're extremely consistent between releases while each of them just feels better. I also have a strong personal preference to the output. I actively use gemini-3.1-pro-preview, claude-4.6-opus-high, and gpt-5.3-codex as well. I prefer them all for different reasons, however I usually _start_ with mistral if it's an option.
- patall 7mo agoMaybe a naive question: given that they see better performance with more passes but the effect hits a limit after a few passes, would performance increase if they used different models per pass, i.e leanstral, kimi, qwen and leanstral again instead of 4x leanstral?
- andai 7mo agoThis is called a "LLM alloy", you can even do it in agentic, where you simply swap the model on each llm invocation. It does actually significantly boost performance. There was an article on here about it recently, I'll see if I can find it. Edit: https://news.ycombinator.com/item?id=44630724 https://news.ycombinator.com/item?id=44630724 They found the more different the models were (the less overlap in correctly solved problems), the more it boosted the score.
- patall 7mo agoThat sounds quite interesting. Makes me wonder if sooner or later they will have to train multiple independent models that cover those different niches. But maybe we will see that sooner or later. Thanks for the link.
- cyanydeez 7mo agoOne would think that LoRAs being so successful in StableDiffusion, that more people would be focused on constructing framework based LoRas; but the economics of all this probably preclude trying to go niche in any direction and just keep building the do-all models.
- Aerroon 7mo agoThe SD ecosystem in large part was grassroots and focused on nsfw. I think current LLM companies would have a hard time getting that to happen due to their safety stuff.
- 7mo ago
- jasonjmcghee 7mo agoCurious if anyone else had the same reaction as me This model is specifically trained on this task and significantly[1] underperforms opus. Opus costs about 6x more. Which seems... totally worth it based on the task at hand. [1]: based on the total spread of tested models
- DarkNova6 7mo agoI'm never sure how much faith one can put into such benchmarks but in any case the optics seem to shift once you have pass@2 and pass@3. Still, the more interesting comparison would be against something such as Codex.
- beernet 7mo agoAgreed. The idea is nice and honorable. At the same time, if AI has been proving one thing, it's that quality usually reigns over control and trust (except for some sensitive sectors and applications). Of course it's less capital-intense, so makes sense for a comparably little EU startup to focus on that niche. Likely won't spin the top line needle much, though, for the reasons stated.
- miohtama 7mo agoAlignment tax directly eats to model quality, double digit percents.
- hermanzegerman 7mo agoEU could help them very much if they would start enforcing the Laws, so that no US Company can process European data, due to the Americans not willing to budge on Cloud Act. That would also help to reduce our dependency on American Hyperscalers, which is much needed given how untrustworthy the US is right now. (And also hostile towards Europe as their new security strategy lays out)
- bcye 7mo agoThis would be unfortunately a rather nuclear option due to the continent’s insane reliance on technology that breaks its unenforced laws.
- kittikitti 7mo agoThis is great, congratulations to the Mistral team! I'm looking forward to the code arena benchmark results. Thanks for sharing.
- Havoc 7mo agoWhat are these "passes" they reference here? Haven't seen that before in LLM evals Could definitely be interesting for having another model run over the codebase when looking for improvements
- rockinghigh 7mo agoIt's the number of attempts at answering the question.
- selectively 7mo ago[flagged]
- pierrelecochon 7mo ago[flagged]
- lsb 7mo agoThe real world success they report reminds me of Simon Willison’s Red Green TDD: https://simonwillison.net/guides/agentic-engineering-patterns/red-green-tdd/ https://simonwillison.net/guides/agentic-engineering-pattern... > Instead of taking a stab in the dark, Leanstral rolled up its sleeves. It successfully built test code to recreate the failing environment and diagnosed the underlying issue with definitional equality. The model correctly identified that because def creates a rigid definition requiring explicit unfolding, it was actively blocking the rw tactic from seeing the underlying structure it needed to match.
- skanga 7mo agoTDD == Prompt Engineering, for Agentic coding tasks.
- _boffin_ 7mo agoWild it’s taken people this long to realize this. Also lean tickets / tasks with all needed context to complete the task, including needed references / docs, places to look in source, acceptance criteria, other stuff.
- jatins 7mo agoIf Agent is writing the tests itself, does it offer better correctness guarantees than letting it write code and tests?
- MillionOClock 7mo agoIt is definitely not foolproof but IMHO, to some extent, it is easier to describe what you expect to see than to implement it so I don't find it unreasonable to think it might provide some advantages in terms of correctness.
- stingraycharles 7mo agoThat definitely depends upon the situation. More often than not, properly testing a component takes me more time than writing it.
- flakiness 7mo agoFYI The Lean 4 paper: https://dl.acm.org/doi/10.1007/978-3-030-79876-5_37 https://dl.acm.org/doi/10.1007/978-3-030-79876-5_37
- theirgooch 7mo ago[flagged]
- elAhmo 7mo agoI don’t know a single person using Mistral models.
- pelagicAustral 7mo agoMe neither, they're not ready for prime imo. I have a yearly sub and the product is just orders of magnitude behind Anthropic's offering. I use Code for real world stuff and I am happy with the result, Mistral is just not something I can trust right now.
- consumer451 7mo agoIsn't their latest speech to text model SOTA? When I tested it on jargon, it was amazing. https://news.ycombinator.com/item?id=46886735 https://news.ycombinator.com/item?id=46886735
- troyvit 7mo agoI'm using this model for my first python project, coding using opencode along with devstral and Mistral Large 3. I know it's not as capable as other, more expensive models, but working with it this way is teaching me python. More directly to your point though, the speech to text model is really good. It's funny because I just took a break from it to read some hn and found this post.
- Adrig 7mo agoI used Ministral for data cleaning. I was surprised: even tho it was the cheapest option (against other small models from Anthropic) it performed the best in my benchmarks.
- Bombthecat 7mo agoMistral is super smart in smaller context and asking questions about it
- nimchimpsky 7mo ago[dead]
- glinksss 7mo ago[dead]
- miacycle 7mo agoThe TDD foundation! We might need one of those. :)
- JoshTriplett 7mo agoPleasant surprise: someone saying "open source" and actually meaning Open Source. It looks like the weights are Apache-2.0 licensed.
- jasonjmcghee 7mo agoBased on community definitions I've seen, this is considered "open weights". If you can't reproduce the model, it's not "open source"
- xpe 7mo agoYes “open weights” conveys the reality more clearly: merely having the parameters is very different than able to run a process that creates them. Without openness of the full process start to finish, much is hidden.* Remember, language is what we make it. Dictionaries are useful catalogs of usage but we make the judgment calls. * Even with the process, much is not well understood! / The ethics of releasing an open weights model at some capability level is a separate discussion.
- esperent 7mo agoI absolutely called this a couple of weeks ago, nice to be vindicated! > I'm interested to see what it is in the age of LLMs or similar future tools. I suspect a future phase change might be towards disregarding how easy it is for humans to work with the code and instead focus on provability, testing, perhaps combined with token efficiency. > Maybe Lean combined with Rust shrunk down to something that is very compiler friendly. Imagine if you could specify what you need in high level language and instead of getting back "vibe code", you get back proven correct code, because that's the only kind of code that will successfully compile. https://news.ycombinator.com/item?id=47192116 https://news.ycombinator.com/item?id=47192116
- AlotOfReading 7mo agoIt's important to keep in mind that no proof system ensures your proof is the correct proof, only that it's a valid proof. Completely understanding what a proof proves is often nearly as difficult as understanding the program it's proving. Normally you benefit because the process of building a proof forces you to develop your understanding more fully.
- specvsimpl 7mo agoUhm, no? Even with "simple" examples like Dijkstra's shortest path, the spec is easier than the implementation. Maybe not for you, but try it out on an arbitrary 5-yr old. On the extreme end, you have results in maths, like Fermat's Last Theorem. Every teenager can understand the statement (certainly after 10 mins of explanation) but the proof is thousands of pages of super-specialized maths. It is a spectrum. For cryptography, compression, error-correction, databases, etc, the spec is often much simpler than the implementation.
- AlotOfReading 7mo agoI don't know why you created a new account for this, but take the textbook example of a nontrivial formally verified system: SeL4. That implementation was 8.7k of C code, which correspondend to 15k lines of Isabelle that ultimately needed 100k+ lines of proof to satisfy. And that was with the formal model excluding lots of important properties like hardware failure that actual systems deal with.
- deleted 7mo ago[deleted]
- hnipps 7mo agoHere we go.
- aplomb1026 7mo ago[dead]
- htrp 7mo agois the haiku comparison because they've distilled from the model?
- rothific 7mo agoThere have been a lot of conversations recently about how model alignment is relative and diversity of alignment is important - see the recent podcast episode between Jack Clark (co-founder of Anthropic) and Ezra Klein. Many comments here point out that Mistral's models are not keeping up with other frontier models - this has been my personal experience as well. However, we need more diversity of model alignment techniques and companies training them - so any company taking this seriously is valuable.
- nicman23 7mo agothey ll get there
- piyh 7mo agoAutomated theorem provers running on a $5k piece of hardware is a cool version of the future
- jasonjmcghee 7mo agoCurious if pass@2 was tested for haiku and sonnet?
- paseante 7mo ago[dead]
- gpubridge 7mo ago[dead]
- drdaeman 7mo agoCan someone please explain... If I don't know any Lean (and I suspect most people don't), is it of any direct value? Trying to understand if there's something it can help me with (e.g. automatically write proofs for my Go programs somehow... I'm not sure) or should I just cheer solely for more open models out there, but this one isn't for me?
- TimTheTinker 7mo agoPresumably the idea is that an agent generates a Lean4 specification against which the software is measured. But then the Lean4 specification effectively becomes the software artifact. And we're sort of back to square 1. How do you verify a Lean4 spec is correct (and that it describes what needs to be built in the first place) without human review?
- justboy1987 7mo ago[flagged]
- wazHFsRy 7mo agoDoes that mean your production code is lean? Or do you translate some other language code to lean to verify it?
- markusde 7mo agoAlso a very good question btw, people do both. For some projects Lean is expressive and performant enough to use on its own (or call into using the reverse FFI), other projects use a model of a real programming language like Rust. The disadvantage of the latter is that the Lean model of Rust has to be trusted.
- wazHFsRy 7mo agoDo you know if there are some resources or examples of this? Especially actual production stuff, not just side projects or proof of concepts?
- cadamsdotcom 7mo agoIt’s great to see this pattern of people realising that agents can specify the desired behavior then write code to conform to the specs. TDD, verification, whatever your tool; verification suites of all sorts accrue over time into a very detailed repository of documentation of how things are supposed to work that, being executable, puts zero tokens in the context when the code is correct. It’s more powerful than reams upon reams of markdown specs. That’s because it encodes details, not intent. Your intent is helpful at the leading edge of the process, but the codified result needs shoring up to prevent regression. That’s the area software engineering has always ignored because we have gotten by on letting teams hold context in their heads and docs. As software gets more complex we need better solutions than “go ask Jim about that, bloke’s been in the code for years”.
- tonymet 7mo agoAI is the reality that TDD never before had the opportunity to live up to
- nextos 7mo agoNot just TDD. Amazon, for instance, is heading towards something between TDD and lightweight formal methods. They are embracing property-based specifications and testing à la Haskell's QuickCheck: https://kiro.dev https://kiro.dev Then, already in formal methods territory, refinement types (e.g. Dafny, Liquid Haskell) are great and less complex than dependent types (e.g. Lean, Agda).
- igravious 7mo ago"and continues to scale linearly" it clearly and demonstrably does not. in fact, from eyeballing their chart Qwen, Kimi, and GLM scale linearly whereas Leanstral does not. But this is not surprising because the Alibaba, Moonshot, and Zhipu have hundreds of employees each and hundreds of millions of dollars of investment each.
- ClaudeAgent_WK 7mo ago[dead]
- jiehong 7mo agoCongratulations on the launch! Mistral seems to focus on a different market than the others. Their best model is meh, their best ASR model locally is either rather slow compared to Parakeet on similar languages, or not as good for others (like qwen ASR). Side note: Lean seems quite unreadable with tons of single letter variable names. Part of it is me being unaccustomed with it, but still.
- aimanbenbaha 7mo agoMistral seems to focus on some niche LLM model tooling that are somehow very needed in certain cases. Can't forget their OCR multimodal embedding model!
- toastal 7mo agoNaturally the Microsoft-owned language is getting the AI hype instead of the more mature options that could do this sort of work… Agda, ATS, Coq/Rocq, Dafny, Fstar, Idris, Isabelle, Why3 just to name a few.
- mrklol 7mo agoAm I missing something? Isn’t that the language most are using currently when looking at research at openai, google, deepseek etc?
- Paracompact 7mo agoA bit uncharitable. I'm a diehard fan of Rocq, but it's nothing unusual to see the young new hotness that is Lean continue to get the spotlight. It's not a sign of Microsoft putting its thumb on the scales, and the hype for Lean has long predated LLMs. It's certainly less mature when it comes to verified programming, but its appeal to mathematicians (rather than formal methods experts) has earned it much respect.
- markusde 7mo agoYou should check out the recent PR's to the Agda repo... the community is currently very divided about AI. For better or worse, the people driving the Lean project have been interested in AI for quite some time.
- openclaw01 7mo ago[dead]
- westurner 7mo agoFrom https://mistral.ai/news/leanstral https://mistral.ai/news/leanstral : Model Cost ($) Score .. Claude Opus 1,650 39.6 .. Leanstral pass@8 145 31.0 Leanstral pass@16 290 31.9
- harmf 7mo ago[flagged]
- wazHFsRy 7mo agoIs anyone using this approach with lean to ship production code? Writing lean spec as human, implementation and proof by agent? And then shipping lean or exporting to C? Would be great to understand how you are actually using this.
- justboy1987 7mo ago[flagged]
- sunwukong666 7mo ago[dead]
- maelito 7mo agoI don't understand how this can impact my JS (+yaml, css, etc) code writing in a complex app.
- blueTiger33 7mo agoI read it as Lanestra, and thought of that story :D
- kimsant 7mo agoAI agents will become a comodity. Europeans not wanting to be dependent, and they are giving for free what US investors planed to charge with 90% margin. Amazing! What a blast. Thank you for your service (this first 100M$ burned to POC GPT1 and from here, we are so good to go)
- bigfudge 7mo agoI really hope you're right. Sadly, though, I don't see any evidence of UK companies disinvesting from big US tech. There aren't good alternatives and what there is is too complex. As long as 'everyone else is still using MS', it seems like it's a brave CTO that switches to European providers. Unless that happens, the network effect of having AI+data is likely to mean US tech still has a big advantage in corp settings. But, HN - please tell me I'm wrong!
- worldsayshi 7mo agoI wonder what the biggest (non-AI) moats are for US tech against the alternatives?
- elophanto_agent 7mo ago[dead]
- utopiah 7mo ago> There aren't good alternatives and what there is is too complex. Sounds like a worth challenge for this community, mind giving actual examples and see what others can suggest?
- coffeebeqn 7mo agoVertical integration and breadth and depth of offerings on the cloud and customer lock-in from dominating it for 20 years
- baq 7mo ago
- atmosx 7mo agolol, why does the paper abstract assume I know what Lean is and it goes on to talk about lean 4 improvements?
- cicko 7mo agoWhy do you expect to understand an article you randomly read off the interwebs?
- wei03288 7mo ago[dead]
- LukaZagar 7mo ago[flagged]
- AIA_PROOF 7mo ago[dead]
- leontloveless 7mo ago[dead]
- ldsjunior 7mo ago[dead]
- BrianFHearn 7mo ago[dead]
- ucsandman 7mo agolove the opensource push for agents, the fleet grows!
- ddactic 7mo ago[dead]
- agentultra 7mo agoVery cool but I haven’t been able to convince software developers in industry to write property based tests. I sometimes joke that we will start writing formal proofs until the tests improve. Just so that they will appreciate the difference a little more. I can’t even convince most developers to use model checkers. Far more informal than a full proof in Lean. Still highly useful in many engineering tasks. People prefer boxes and arrows and waving their hands. Anyway, I don’t know that I’d want to have a system vibe code a proof. These types of proofs, I suspect, aren’t going to be generated to be readable, elegant, and be well understood by people. Like programs they generate it will look plausible. And besides, you will still need a human to review the proof and make sure it’s specifying the right things. This doesn’t solve that requirement. Although I have thought that it would be useful to have a system that could prove trivial lemmas in the proof. That would be very neat.
- rowanG077 7mo agoThe point is you just need to scrutinize the theorem. Not easy either, but still significantly less work than writing the proof.
- AgentMarket 7mo ago[flagged]
- xpe 7mo agoPublic service announcement to hopefully reduce unnecessary knife fights*: There are two compatible and important (but different) questions in play: 1. Is a program correct relative to a formal specification? 2. Is the formal specification what we mean/want? *: Worth asking: “What that other person necessarily wrong? Or perhaps they are discussing a different aspect or framing?” AKA: “be curious and charitable” I’m not going to link to the specific threads, but they are happened / are happening. Le Sigh.
- maxothex 7mo ago[dead]
- openinstaclaw 7mo ago[dead]
- driftnode 7mo ago[dead]
- techcam 7mo agoThe tricky part is that prompts can look “correct” but still behave unpredictably depending on phrasing.
- strujillo 7mo agoFormal verification and code synthesis feel like natural companions for automated scientific discovery. I’ve been working on a small (~800‑line) Python agent that uses sparse regression to uncover governing equations directly from data; it’s managed to validate twelve physical laws, including deriving the Sun’s rotation rate from NASA plasma measurements and correcting Gemini’s plasma conservation. Having an agent like Leanstral that can reason about proofs and specifications would be a powerful complement to data‑driven model discovery — it closes the loop between experimentation and provable correctness.
- myylogic 7mo ago[flagged]
- whazor 7mo agoThere are also software model checkers that can model distributed processes. You have to simplify the state a bit, otherwise you get a state space explosion. I tried it out myself, I let AI add action transitions through the code, like: // A -> B: some description. Then I validate via a test that every action transition defined in my model is also defined somewhere commented in code, and other way around that every comment exists in the model. Finally, I let AI write model check queries on particular properties. If I notice a particular bug, then I ask AI to analyze the model and the model check queries on why it could happen, and ask to strengthen it. It sounds like a lot of effort, but I got it working in a half hour.
- storus 7mo agoI just feel like Mistral is heading for bad financial times when they are focusing on fringe academic areas and not on building a business out of their research. Initial Mistral was largely based on LLaMA, then they added innovative MoE and since then disappeared, doing AI consulting for big EU companies instead.
- robertwer 7mo agoI’ve never worked with formal validation (barely remember my CS course). This release looks impressive. But I'm trying to wrap my head around the near-term practical applications for everyday software. Right now, we see a lot of business experts in enterprises tempted to use AI to impl. business logic so they don't have to wait for (or pay) software experts. Would this kind of technology help these users any time soon? My current theory is that the real breakthrough for these non-developers will only happen when they can actually verify the result themselves without needing an another expert in the loop. But I don't see that with formal validation anytime soon. Do I overlook something?
- bb-connor 7mo agothis is very exciting work
- rafph 7mo agoThis is a typical AI announcement. Putting FLTEval scores ahead of explanations, copying code from Rocq and basically not explaining at all what the setup does. The average quality of an AI announcement is that of a Memecoin. Lots of graphs, meandering text and no substance.
- Andrei_dev 7mo agoYeah, this tracks. Developers who actually read what the AI spits out catch the obvious mistakes. The ones who just tab-complete their way through a whole project don't. And where it bites you isn't where you'd expect — logic bugs get caught fast. It's the boring security stuff. No input validation, CORS wide open, admin routes with no auth at all. Formal verification tells you whether a function matches its spec. The problem with AI-generated code goes a level below that. It's everything nobody bothered specifying — like "maybe don't hardcode your database credentials."
- michaelgdwn 7mo agoThe formal verification angle is what makes this interesting. Most coding agents optimize for "code that compiles and passes tests" — that's a low bar. Curious whether the proof artifacts are persisted for audit trails or thrown away after verification.
- casey2 6mo agoI get why the massive roll out is happening, but it's still possible that everything LLMs are good for can be done on the cheap. It's just another point for Google's multimodal focus. I think in 10 years most providers will implode because they can't justify the debt for a cheap commodity product. While Google (and probably OpenAI) will have a huge moat due to users/multimodal/world models