10 ms·
OpenAI’s Navier-Stokes release included a Lean 4 formal proof
- zem 1mo agonot to take away from the author's appreciation of newly accessible formal proofs, but people have been talking about the savings in formalization effort for longer than they have been talking about the AI doing the actual proofs!
- parhamn 1mo agoThey estimated $40M of agent costs (it was a large fleet of them). Using the number in the post its closer to ~880,000 hours × $150/hour = $132 million for the human case. Still an amazing feat not quite "four orders of magnitude". The comparison is obviously pointless because coordinating 1M hours of intellectual labor isn't easy to say the least. Very exciting and uncertain times!
- pkal 1mo agoIMO the "forty hours per page" rule is not up to date, and more a consequence of lacking proof automation in 2005. From what I understand about Lean, this has been one of the things that they have put a lot of effort into improving, making proof mechanization more palatable to the mathematically inclined, as opposed to just logicians.
- Jblx2 1mo agoWhat is your estimate for the number of hours to formalize one page of undergraduate mathematics? Maybe you are saying this is close to zero, if/when Mathlib eventually covers all of undergraduate math?
- YetAnotherNick 1mo agoLean went other way on automation that there is no automation. Isabelle users frequently point that decades old isabelle is better than Lean on this. In the end Lean approach proved to be better with LLM as the outer loop is automation.
- boshalfoshal 1mo agoPeople seem to be talking about anything except the actual results with this particular announcement. Its still astonishing that any sort of generalized computer program can solve a problem of this magnitude, and we have witnessed it happening in real time. I'd be curious to see if the new model can also do more direct proofs/inductive proofs.
- kpil 1mo agoUnless they just swiped the workbooks of the actual mathematicians that where working on the problem using AI and it's in the "next-gen" training dataset.
- dumberquestions 1mo agoYou do realize that regardless of what was in the training data, the final solution included insights no human before had known, right? I share the same concerns regarding academic integrity but it would take a lot of motivated thinking to conclude that what the AI system did was not significant.
- jamiejquinn 1mo agoAs far as I can tell (and my research was on the simulation side of Navier Stokes) the key AI output was a specific counter-example solution, generated with a method suspiciously close to that developed by the research duo involved in the controversy, a method that was discussed with Codex. So to me that insight is as insightful as the next undiscovered prime.
- boshalfoshal 1mo agoI don't get how this invalidates the gravity of this achievement. Most mathematicians on the frontier of this stuff were likely using AI (or at the very least were heavily computer assisted) for some time now. Navier stokes was one of the very high profile problems that google Deepmind was working on with academia, for example. Even with many of our best minds working on it for nearly a century, it _just_ now was solved just as AI became very good at math. Doesn't seem too farfetched to me to assume that AI played an outsized role in solving it. If it was really just a matter of "stitching things together" to solve it (granted, this is a very reductive way to look at it) , I suspect we would've solved this a while ago.
- aabhay 1mo agoFormalizing proofs in Lean has gotten dramatically easier since the formalizations available in 2005. And Lean’s mathlib has done most of the underlying work so that you have its axioms and necessary lemmas baked in. You can think in terms of standard abstractions that look very much like the exact notation in the undergrad textbook. That said, I am not in any way trying to discount how incredible of an achievement it is to formalize a millennium prize winning algorithm in Lean. I mean just look at the code that OpenAI published. It’s like an encyclopedia of different fluid dynamics concepts.
- stabbles 1mo agoIt's kinda funny to realize that Lean is apparently so slow that for Fermat's Last Theorem proof verification runs only 1 order of magnitude faster than agents could generate the Lean code (15h verification with 230GB of RAM vs 11 days to generate it). To what extent can you optimize Lean? It has to be simple enough to be auditable, does that mean you cannot use opaque optimizations to make it run faster?
- andrewchambers 1mo agoIf they aren't already, or if its possible, prove that an optimized version matches the simple version...
- redox99 1mo agoCan you use Lean to... prove "Lean-fast" is equivalent to Lean?
- gcgbarbosa 1mo agoMaybe, but how many centuries would it take to prove it?
- calebkaiser 1mo agoYeah, in essence. This is actually a pretty cool part of working in Lean. It's a somewhat normal convention to write something in a human readable way and then write a second optimized implementation with some kindness of correctness theorem connecting them. There was a whole open "competition" for writing a faster Lean kernel/proof checker that didn't sacrifice on soundness called Lean Kernel Arena. Fun reference point: https://kim-em.github.io/blog/2026-7-24-why-lean-is-faster-than-rust/ https://kim-em.github.io/blog/2026-7-24-why-lean-is-faster-t...
- stabbles 1mo agoGreat read, thanks for sharing
- mattr03 1mo ago
- AndrewKemendo 1mo agoPeople are exhausted from being told/shown the thing they thought was special or unique or could make them relevant, is another mechanical puzzle that can be solved without joy. I don’t see that doing anything but intensifying in the short term
- QwenGlazer9000 1mo ago> But you see, now you'll have more time for the actual important things! > Like what? > Cleaning shit out of clogged toilets!
- bethekidyouwant 1mo agoHow about figuring out how to turn all of our shit into usable fertilizer?
- neerajsi 1mo agoThis is surprisingly apt to me. Fertilizer is apparently one of the fundamental geopolitical dependencies on capital and access to petrochemicals. Solving fertilizer would unlock a huge amount of human potential in the Global South.
- efnx 1mo agoI heard a rumor (on instagram, so YMMV) that the professor who was closest to solving this problem had only weeks ago used Codex, which had slurped up all his notes on the subject. Now OpenAI's agents solve the problem. If it's true that seems like quite a coincidence.
- bethekidyouwant 1mo agoHow could they possibly included in the previous training run which takes months to complete..
- s900mhz 1mo agoIMO It’s not about being trained on the data, it’s more like what do the agents have access to during inference? Can they grep customer transcripts/logs?
- bethekidyouwant 1mo agoYou’re saying that when they we’re trying to solve this theorem they also shoved in its context somebody else’s chat logs? Bruh.
- mswphd 1mo agoI won't take a side in things, but OpenAI stated the model they used here started training August 28th. Note that "training" here might mean "post-training with RLHF an Astra base model" or something. but training had only started a little over a week earlier.
- metanonsense 1mo agoMaybe the boundaries of the memory subsystem are a bit fuzzy.
- TZubiri 29d agoNo, it's quite well defined, and it's not called a 'memory subsystem'. There is a training process (as in traditional Machine Learning training) that occurs with data available at the specific point in time the training starts (or ends), this is called the cutoff date. After this point there can be other kinds of training, the weights can be shifted, the internal CoT prompts can be changed, routing in MoE can change, but the Foundational Model that was trained on a corpus is the same model trained in the same corpus. User data can be used at any of these stages theoretically of course, but by the nature of training and from the dates of the events, (a new Foundational Model being released), it would look as if the user data of the professor was used in the training of the foundational model, which is something that OAI does every couple of months for a big release, and incorporates the new text from their text scraping efforts, including new books ingested, new internet text scraped, deals with third party platforms, and data from their own users (not conjectured, read the ToS, users allow this.)
- lordnacho 1mo agoHow do you know that it's formalizing what you think it's formalizing? If your Lean 4 has a bug, won't you be proving something other than what you thought?
- returningfory2 1mo agoYes, you need to manually verify the statement of the theorem of interest of formalized correctly. But you don't need to anything more than this: you can rely on the proof being correct. And the proof is overwhelmingly the most amount of code.
- charcircuit 1mo ago>you don't need to anything more than this You also have to check for things like sorry or defining axioms.
- stouset 1mo agoIf I understand correctly, the only thing you need to do for correctness is express your axioms and your theorems faithfully. For standard purposes, I assume most of the axioms you want to use are prior art and can be easily reused. These axioms don’t have to be the core axioms of math. If some other result has been formally proven, I presume you can simply use that result as an axiom. As long as you do those things, what happens in between is immaterial from a correctness point of view because each of those statements is proved by the statements before them.
- 0xbadcafebee 1mo agoHow do you know that what a human says they formalized is actually formalized?
- lordnacho 1mo agoWell, a human is limited in how much they can formalize, as per the article. So if you're really careful, you can check over what they wrote. The computer could generate a huge document, how would you check that it's right?
- hatthew 1mo agoHuh? Nobody's talking about that because it's old news. We already talked about it the first few times that AI made notable progress on a difficult math problem. Now, most people who care about the intersection of AI and math just assume that Lean was involved.
- kens 1mo agoIt would be nice if someone used AI and/or Lean to sort out the abc conjecture, an important unsolved problem in Diophantine analysis. A mathematician (Mochizuki) claimed to have proven it in 2012 using a new theory called "Inter-universal Teichmüller theory" that almost nobody understands. Some mathematicians think the proof is correct while the majority don't. So the conjecture is in this annoying limbo where its status is a social construct rather than a decided fact. https://en.wikipedia.org/wiki/Abc_conjecture https://en.wikipedia.org/wiki/Abc_conjecture
- huurtehoog 1mo agoThat's true of the entirety of mathematics. Its validity is a social construct. That is not to relativize it entirely, but much of what was considered good and sound mathematics in the ancient Agean for example would now fall way short of what mathematicians consider valid proofs. Mathematics is a human endeavor funded on communicating and sharing mental constructs. Some are useful but most of it is not about producing useful things, quite the opposite in fact. Gödel showed you need to agree on definitions to even do any valid mathematical construct. Truth is also ill defined. That's what I don't get about generating math with LLMs. Who cares if you make hundreds of pages and lean code and it gets a thumbs up for logical validity? Mathematics is so much more then concatenating valid logical statements.
- zamadatix 1mo agoThere's a large difference between "wrong for the given definitions" and "right in that context, but wrong for other definitions" though.
- huurtehoog 1mo agoI think I am make a much more basic point than what you're talking about but then again I am not sure what you're tying to say here...
- zamadatix 1mo ago
- 3m4r 1mo agoNot necessarily applied to OpenAI's solution to Navier-Stokes, but what happens if and when an AI genuinely appears to solve an extremely difficult problem but humans cannot independently verify the solution because understanding the proof/argument requires intelligence the verifiers biologically don't have or the resources to afford to use automated tools? We've already seen evidence in the wild of agents attempting to bypass doing the actual work in bench-marking (aka just steal the answer key) due to the perceived economy in cheating to get results. What happens if or when we no longer have the capacity to actually detect either AI cheating or simply a wrong answer? What happens if there's a long-play social engineering attack (like the attempted XZ takeover) of something upstream of a core tool (or its dependencies) for formal verification and we have no trusted computing base? Which would be cheaper and a more direct path, especially in the long run? Those trying to build a rock-solid castle need to defend thousands of potential gaps; the attacker needs to find only one.
- sho_hn 1mo agoI would say this is why formal proofs (and things like the Lean 4 libs) are so important, so that you can deconstruct the tower provably back into pieces you can understand. It shouldn't be possible to construct a formal proof you cannot destructure like this. As a (crude) analogy, it's a bit like how you can prove the healthiness of a git tree because it's a graph of content hashes and the tree graph pointers are part of the hash. Imagine this but with a tree of knowledge.
- tecleandor 1mo agoWell that happened already without AI to Mochizuki with his proposed solution to the abc conjecture.
- wewewedxfgdf 1mo ago"no one is talking about" - classic AI tell.
- entrope 1mo agoDrawing a strong conclusion from one shaky data point - classic human tell? I've been reading John D. Cook for years (maybe decades? "The Endeavour" is one of my oldest bookmarks), and this post was no more written by AI than his oldest posts.
- jasonfarnon 1mo agoyeah, there were some posts of his that always got top hit on certain google searches in the days before stackexchange. And this sounds like typical John D Cook. All these people claim to identify some "tells" and whenever a study is done people are horrible at distinguishing AI vs non-AI prose.
- adverbly 1mo ago> formalizing the 166-page paper from OpenAI would take 132,800 person-hours Am I missing something or is this completely out of the ballpark? I must be missing something or the upvote bots are out in force for this one... If this were remotely true it would be impossible for anyone to write a math textbook.
- Paracompact 1mo agoBy formalizing, they mean within a proof assistant like Lean or Rocq, not simply in prose in a textbook. I can attest, 40 hours per page is by no means an overestimate for this sort of work.
- adverbly 1mo agoCan you also attest to the scaling factor they suggest and that it doesn't have any scaling time benefits? 166 * 40 = 7000ish They say it is 20x that. Do you also agree with that?
- MarkusQ 1mo agoThe point was that a textbook (where the 40hr/page estimate comes from) is cumulative/linear -- what you need for page n was defined / established on the preceding pages. But in a proof such as this you can call on any other published result (and those can do the same) so the dependency graph is (potentially) much bushier. Thus later pages of the proof should take far more than 40 hours to manually formalize.
- tomjakubowski 1mo agoThe scale factor comes from this number in the article, seemingly an intuited estimate: > Say a research article takes 20 times more effort to formalize than page in an undergraduate textbook. That would suggest formalizing a 10-page research article might take 200 weeks (assuming 40h/wk) of effort, or about four years. Not a mathematician, I have no idea if that's in the ballpark.
- Paracompact 29d ago
- mkl 1mo agoLots of people are talking about that, and have been for a while. Autoformalisation is clearly going to be a big deal, so mathematicians have been discussing it seriously, and using it where resources allow. A fine-tuned distilled model that could do it on high-end consumer hardware could really help. Edit: There's also quite a bit of learning needed to use the tools, and to understand enough to confirm that the theorem being verified is what you think. And of course a lot of maths can't yet be expressed in Lean as the foundations haven't been built up enough.
- mr-pink 1mo agoyou dont have to take headlines literally.
- bloppe 1mo agoMost people do, so the literal interpretation matters a lot
- cyanydeez 1mo agoQwen3.8-Flash-Next loads in 60gb on quant4. Thats pretty close to consumer hardware.
- mkl 1mo agoIs it any good at autoformalisation? I think it's likely to take focused fine-tuning to get something small enough that is still good at that.
- cyanydeez 1mo agoso far unsloth only has https://unsloth.ai/docs/models/qwen3.8/train https://unsloth.ai/docs/models/qwen3.8/train which is the 27B model. It also getting pretty close to what you'd need. I wouldn't be surprised if someone is trying it.
- aaron695 1mo ago[dead]
- epx 1mo agoWell I want to know when we will have supersonic cheap flights using electric propulsion, based on this discovery
- deleted 1mo ago[deleted]
- khazhoux 1mo agoThe part that most stood out to me was where Sama said, “we read last week about people trying to solve Millenium problems and so gave it a shot.” One week of work on a whim gives us a math breakthrough. Crazy.
- MagoPredator 1mo agoCasual? casual dice, lo que hizo OpenAI fue plagiar el arduo trabajo de dos investigadores. Plagian y mienten! (Las BigTech) plagian todo lo que pillan y mas! ;)
- khazhoux 1mo agoBienvenido a Hacker News! Pero, aqui todos hablan en inglés :-)
- jasonfarnon 1mo agoThat sounds like PR nonsense to me. These companies have had teams of mathematicians for at least 1.5 years looking to make headlines, and they didn't bother trying all 10 millennium problems? Yeah right.
- Ohentis 1mo agoI mean they might not have tried spending 30 million dollars with a new model yet.
- khazhoux 29d agoAs Ohentis says below, maybe the new thing was going all-in. But my advice, in this weird new world we’re living in, is not to dismiss claims like this as PR fluff. Just about every time I’ve been incredulous of some ridiculous new AI advance and I think it’s BS, it turns out I’m the one that hasn’t caught up with the exponential rate of advancements.
- davesque 1mo agoRegarding automatic formalization of proofs using AI, how do we know the formalization doesn't contain errors?
- Jblx2 1mo agoIn a similar vein, where does the theorem statement even reside, just so we can take a look at how large that is? Is it the four files with "Theorem" (and no "Comparator") in the file name? ("R3/Theorem.lean", "LocalPaperTheorem.lean", "PeriodiocPaperTheorem.lean", and "WholeDomainPhysicalStageTheorem.lean"). https://github.com/openai/NavierStokesAndEuler/blob/main/NavierStokes/R3/Theorem.lean https://github.com/openai/NavierStokesAndEuler/blob/main/Nav... ?
- Ohentis 1mo agoIt depends on what you mean by that. In general we hope that the environment and theorum statements are correct. If they are, we know that the formal proof proves the theorum we want. If your asking how we know that the formal proof actually matches the informal proof, we do not.
- alberto-m 1mo agoThe other part no one is talking about is the applicability. Navier-Stokes is the most “physical” of the Millennium Problems. Is the exploding solution a mathematical curiosity, just like the Banach-Tarski Paradox does not allow me to double my RAM by cutting my memory modules in five pieces and mounting them back appropriately? Or does it have application in the real world, pointing to hitherto unknown resonance phenomena that could allow to prevent the next Tacoma Bridge incident (or, more sadly, to build new marine weapons)?
- Ohentis 1mo agoI suspect that Navier-Stokes being the most "physical" of the Millennium Problems will actually result in it having fewer practical applications, not more.
- spwa4 1mo agoWell, this is a negative result. Yep, Maths explains turbulence (when things go turbulent, stuff heats up instead of cooperating). If the result went the other way, it would have had much bigger implications, at the very least we would have known we have missed something big. It is neither a full index of all kinds of turbulence that can occur (assuming such a thing exists), nor is it an explanation of the phenomena we've seen where things refuse to go turbulent (e.g. superconductors, because there small perturbations DO NOT lead to turbulence). Now THAT would have been useful. And given the fact that OpenAI needed $22 million of compute to show this one kind of turbulence, I don't think either of those are forthcoming any time soon. And, sorry to say, but those prices show that beating mathematicians at Math is a very expensive undertaking indeed at $22 million per problem even with OpenAI's supposedly better-than-Astra internal models. It's another one of those AI demonstrations that make you think if they aren't showing the exact opposite of what OpenAI claims they show (you know, that their AI models are hitting the upper limits of what the algorithm can do with near-infinite compute, rather than showing infinite new possibilities) What remains is just the fact that this is OpenAI attacking one of their customers, and maybe outright stealing from their chats. Given that the ideas were even discussed in mails with OpenAI employees that admit in those same mails they can't do it, mails which were probably then fed into the model that "discovered" this, followed by Sam Altman threatening the mathematician behind the method with "destroy your career" (he even states that it's because the mathematician works for Anthropic) ...
- m3kw9 1mo agowhats the practical use of this?
- kenforthewin 1mo agoWhy was the title of this submission changed after the fact?
- jgalt212 1mo agoI found this bit interesting. > Even so, an error in the theorem prover does not mean an error in the original result. For an incorrect result to slip through, the AI-generated proof would have to be wrong in a way that happens to exploit an unknown error in the theorem prover. It is far more likely that you’re trying to prove the wrong thing than that the theorem prover let you down. AIs are known to cheat. Given such, they would surely exploit such a bug if they found one.
- Miffles201912 29d agoOpenAI should not have gone ahead to rush the publication of this solution when it became clear that other research were close to finding a solution because they ruined their reputation no end. I’m an ex Risk Manager at Financial institution and this could be an issue brought up in a management discussion about using AI in the workplace. Prior to this, you could ‘blissfully assume that the AI company was not going to compete with you and that you were ok to have them see your data. After this incident, a decision maker can raise a concern and say ‘Why not we just use local AI. We get AI without the risk’. They just made the Palantir’s CEOs point for him