6 ms·
First Proof
https://www.nytimes.com/2026/02/07/science/mathematics-ai-proof-hairer.html https://www.nytimes.com/2026/02/07/science/mathematics-ai-pr...
https://www.scientificamerican.com/article/first-proof-is-ais-toughest-math-test-yet-the-results-are-mixed/ https://www.scientificamerican.com/article/first-proof-is-ai...
- happa 8mo agoFebruary 13th is a pretty close deadline. They should at least have given a month.
- blenderob 8mo agoFebruary 13 seems right to me. I mean it's not like LLMs need to manually write out a 10 page proof. But a longer deadline can give human mathematicians time to solve the problem and write out a proof. A close deadline advantages the LLM and disadvantages humans which should be the goal if we want to see if LLMs are able to solve these.
- mlpoknbji 8mo agoThis is a very interesting contribution to the AI/math space. I hope it can be seen by nonmathematicians interested in this. The mathematicians involved are quite well known (Martin Hairer is a Fields medalist). See https://www.reddit.com/r/math/comments/1qx77l7/a_new_ai_mathematics_assessment_that_was_designed/ https://www.reddit.com/r/math/comments/1qx77l7/a_new_ai_math... for some discussions.
- samasblack 8mo agohttps://1stproof.org/#about https://1stproof.org/#about
- baal80spam 8mo agoI'll patiently wait for the "goalpost moving olympics" after this is published.
- blenderob 8mo agoThe goalposts have been on wheels basically since the field was born. Look up "AI effect". I've stopped caring what HN comments have to say about whether something is or isn't AI. If its useful to me, I'm gonna use it.
- blenderob 8mo agoCan someone explain how this would work? > the answers are known to the authors of the questions but will remain encrypted for a short time. Ok. But humans may be able to solve the problems too. What prevents Anthropic or OpenAI from hiring mathematicians, have them write the proof and pass it off as LLM written? I'm not saying that's what they'll do. But shouldn't the paper say something about how they're going to validate that this doesn't happen? Honest question here. Not trying to start a flame here. Honestly confused how this is going to test what it wants to test. Or maybe I'm just plain confused. Someone help me understand this?
- yorwba 8mo agoThis is not a benchmark. They just want to give people the opportunity to try their hand at solving novel questions with AI and see what happens. If an AI company pulls a solution out of their hat that cannot be replicated with the products they make available to ordinary people, that's hardly worth bragging about and in any case it's not the point of the exercise.
- cocoto 8mo agoThey could solve the problems and train the next models with the answers, as such the future models could “solve” theses.
- fph 8mo agoThe authors mention that before publications they tested these questions on Gemini and GPT, so they have been available to the two biggest players already; they have a head start.
- data_maan 8mo agoLooks like very sloppy research.
- pickleRick243 8mo agoI don't think it's that serious...it's an interesting experiment that assumes people will take it in good faith. The idea is also of course to attach the transcript log and how you prompted the LLM so that anyone can attempt to reproduce if they wish.
- falloutx 8mo agoAnything special about these questions? Are they unsolved by humans. I am not working in mathematics research so its hard to tell the importance.
- jsnell 8mo agoThe abstract of the article is very short, and seems pretty clear to both of your questions. This is what is special about them: > a set of ten math questions which have arisen naturally in the research process of the authors. The questions had not been shared publicly until now; I.e. these are problems of some practical interest, not just performative/competitive maths. And this is what is know about the solutions: > the answers are known to the authors of the questions but will remain encrypted for a short time. I.e. a solution is known, but is guaranteed to not be in the training set for any AI.
- blenderob 8mo ago> I.e. a solution is known, but is guaranteed to not be in the training set for any AI. Not a mathematician and obviously you guys understand this better than I do. One thing I can't understand is how they're going to judge if a solution was AI written or human written. I mean, a human could also potentially solve the problem and pass it off as AI? You might say why would a human want to do that? Normal mathematicians might not want to do that. But mathematicians hired by Anthropic or OpenAI might want to do that to pass it off as AI achievements?
- teraflop 8mo agoWell, I think the paper answers that too. These problems are intended as a tool for honest researchers to use for exploring the capabilities of current AI models, in a reasonably fair way. They're specifically not intended as a rigorous benchmark to be treated adversarially. Of course a math expert could solve the problems themselves and lie by saying that an AI model did it. In the same way, somebody with enough money could secretly film a movie and then claim that it was made by AI. That's outside the scope of what this paper is trying to address. The point is not to score models based on how many of the problems they can solve. The point is to look at the models' responses and see how good they are at tackling the problem. And that's why the authors say that ideally, people solving these problems with AI would post complete chat transcripts (or the equivalent) so that readers can assess how much of the intellectual contribution actually came from AI.
- _alternator_ 8mo agoThese are very serious research level math questions. They are not “Erdős style” questions; they look more like problems or lemmas that I encountered while doing my PhD. Things that don’t make it into the papers but were part of an interesting diversion along the way. It seems likely that PhD students in the subfields of the authors are capable of solving these problems. What makes them interesting is that they seem to require fairly high research level context to really make progress. It’s a test of whether the LLMs can really synthesize results from knowledge that require a human several years of postgraduate preparation in a specific research area.
- clickety_clack 8mo agoSo these are like those problems that are “left for the reader”?
- Jaxan 8mo agoNot necessarily. Even the statements may not appear in the final paper. The questions arose during research, and understanding them was needed for the authors to progress, but maybe not needed for the goal in mind.
- jasonfarnon 8mo agoNo, results in a paper are identified to be "left for the reader" because they are thought to be straightforward to the paper's audience. These are chosen because they are novel. I didn't see any reason to think they are easier than the main results, just maybe not of as much interest.
- data_maan 8mo agoVery serious for mathematicians - not for ML researchers. If the paper would not have had the AI spin, would those 10 questions still have been interesting? It seems to me that we have here a paper that is solely interesting because of the AI spin -- while at the same time this AI spin is really poorly executed from the point of AI research, where this should be a blog post at most, not an arXiv preprint.
- richard_chase 8mo agoInteresting questions. I think I'll attempt #7.
- gre 8mo agoTried all ten with claude, then had codex take a loook at the work -- codex thinks number 7 has the lowest chance of being correct, a 1 out of 10 rating. None of them were higher than 7/10 chance of being right so far as done by claude opus 4.6 and evaluated by codex 5.3 highest. Not going to spend too many more tokens on this.
- pickleRick243 8mo agoI don't think either of these are the best choices for this. Chatgpt 5.2 pro and gemini 3 pro deep thinking I believe are the strongest LLMs at "pure thought", i.e. things like mathematical reasoning.
- The_Gray 8mo agoAny chance you're willing to share the links/outputs?
- Aressplink 8mo ago[flagged]
- Syzygies 8mo agoI'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of life because we're a selfish species. In contrast Niki Lauda trusted fellow Formula 1 race car driver James Hunt to race centimeters apart. Some people want AI to help them drive that well. They'll have great jobs as AI evolves. Gary Kasparov pioneered "freestyle" chess tournaments after his defeat by Big Blue, where the best human players were paired with computers, coining the "centaur" model of human-machine cooperation. This is frequently cited in the finance literature, where it is recognized that AI-guided human judgement can out-perform either humans or machines. Any math professor knows how to help graduate students confidently complete a PhD thesis, or how to humiliate students in an oral exam. It’s a choice. To accomplish more work than one can complete alone, choose the former. This is the arc of human evolution: we develop tools to enhance our abilities. We meld with an abacus or a slide rule, and it makes us smarter. We learn to anticipate computations, like we’re playing a musical instrument in our heads. Or we pull out a calculator that makes us dumber. The role we see for our tools matters. Programmers who actually write better code using AI know this. These HN threads are filled with despair over the poor quality of vibe coding. At the same time, Anthropic is successfully coding Claude using Claude.
- wizzwizz4 8mo agoThat centaurs can outperform humans or AI systems alone is a weaker claim than "these particular AI systems have the required properties to be useful for that". Chess engines consistently produce strong lines, and can play entire games without human assistance: using one does not feel like gambling, even if occasionally you can spot a line it can't. LLMs catastrophically fail at iterated tasks unless they're closely supervised, and using LLMs does feel like gambling. I think you're overgeneralising. There is definitely a gap in academic tooling, where an "association engine" would be very useful for a variety of fields (and for encouraging cross-pollination of ideas between fields), but I don't think LLMs are anywhere near the frontier of what can be accomplished with a given amount of computing power. I would expect simpler algorithms operating over more explicit ontologies to be much more useful. (The main issue is that people haven't made those yet, whereas people have made LLMs.) That said, there's still a lot of credit due to the unreasonable effectiveness of literature searches: it only usually takes me 10 minutes a day for a couple of days to find the appropriate jargon, at which point I gain access to more papers than I know what to do with. LLM sessions that substitute for literature review tend to take more than 20 minutes: the main advantage is that people actually engage with (addictive, gambling-like) LLMs in a way that they don't with (boring, database-like) literature searches. I think developing the habit of "I'm at a loose end, so I'll idly type queries into my literature search engine" would produce much better outcomes than developing the habit of "I'm at a loose end, so I'll idly type queries into ChatGPT", and that's despite the state-of-the-art of literature search engines being extremely naïve, compared to what we can accomplish with modern technology.
- data_maan 8mo agoAs mathematically interesting the 10 questions are that the paper presents, the paper is --sorry for the harsh language-- garbage from the point of view of benchmarking and ML research: Just 10 question, few descriptive statistics, no interesting points other than "can LLMs solve these uncontaminated questions", no long bench of LLMs that were evaluated. The field of AI4Math has so many benchmarks that are well executed -- based of the related work section it seems the authors are bit familiar with AI4Math at all. My belief is that this paper is even being discussed solely because a Fields Medalist, Martin Hairer, is on it.
- bawolff 8mo agoPaper not about benchmarking or ML research is bad from the perspective of benchmarking. Not exactly a shocker. The authors themselves literally state: "Unlike other proposed math research benchmarks (see Section 3), our question list should not be considered a benchmark in its current form"
- data_maan 8mo agoOn the website https://1stproof.org/#about https://1stproof.org/#about they claim: "This project represents our preliminary efforts to develop an objective and realistic methodology for assessing the capabilities of AI systems to autonomously solve research-level math questions." Sounds to me to be a benchmark in all but a name. And they failed pretty terribly at achieving what they set out to do.
- bwfan123 8mo ago> And they failed pretty terribly at achieving what they set out to do. Why the angst ? If the ai can autonomously solve these problems, isnt that a huge step forward for the field.
- data_maan 8mo agoIt's not angst. It's intense frustration that they 1) are not doing the science correctly, and 2) that others (e.g. FrontierMath) already did everything they claim to be doing, so we won't learn anything new here, but somehow 1stproof get all the credit.
- hiq 8mo agoI'm realizing I don't know if it's currently harder for an LLM to: * come up with a formal proof that checks out according to a theorem prover * come up with a classical proof that's valid at a high-level, with roughly the same correctness as human-written papers Is this known?
- pama 8mo agoThe advantage of the formal proof is that the LLM in a loop can know that it failed and keep trying.
- Western0 8mo agoNo, this is not a proof because not using Mizar ;-) https://mizar.uwb.edu.pl/ https://mizar.uwb.edu.pl/
- LegionMammal978 8mo agoWould something be a proof in that sense even if it did use Mizar? As far as I can tell, Mizar has no complete reference for its language semantics, except for the single closed-source implementation. In general, information about the system itself (outside of the library) seems very scarce, or at least scarcely advertised.
- robinzfc 8mo agoMizar source was "available upon request" for maybe 30-40 years. It got completely open-sourced under GPL some 3 years ago (maybe earlier, not sure), see [1], also [2] and [3] about an alternative implementation in Rust. Mizar is indeed "scarcely advertised", but all the information is publicly available, who wants to know knows. As for Mizar semantics, see for example [4]. [1] https://github.com/MizarProject/system https://github.com/MizarProject/system [2] https://github.com/digama0/mizar-rs https://github.com/digama0/mizar-rs [3] https://arxiv.org/pdf/2304.08391v2 https://arxiv.org/pdf/2304.08391v2 [4] https://link.springer.com/article/10.1007/s10817-018-9479-z https://link.springer.com/article/10.1007/s10817-018-9479-z
- LegionMammal978 8mo agoThank you for that information, all I could find on the website was that "The source code of the Mizar verifier and accompanying tools is available to the members of SUM" [0], which of course does not reflect the newer status quo. > Mizar is indeed "scarcely advertised", but all the information is publicly available, who wants to know knows. Yes, there indeed seems to be a good bit of information available, especially about the library and its articles. But some parts seem to be scattered about, unless you already know where to look, or know someone who knows. Perhaps it's a matter of taste. (For comparison, I've recently been dabbling a fair bit with Metamath: it's not really advertised outside of its small circle these days, but the website does a good job at introducing the system, while also offering a complete reference in the form of the Metamath book. From there, the primary challenges to a new user are the fiddly tooling, the cryptic labeling scheme, and the puzzling DV conditions.) [0] https://mizar.uwb.edu.pl/system/ https://mizar.uwb.edu.pl/system/
- rvz 8mo ago> Conflicts of interest. No funding was received for the design or implementation of this project. None of the authors of this report was employed by or consulted with AI companies during the project, nor will they do so while contributing to it As it should. Good. This is a totally independent test not conducted or collaborated by any of the AI companies or employees so that no bias is introduced at all[0]. [0] Unless the researchers are not disclosing if they have any ownership of shares in private AI companies.
- mlpoknbji 8mo agoNYTimes: https://www.nytimes.com/2026/02/07/science/mathematics-ai-proof-hairer.html https://www.nytimes.com/2026/02/07/science/mathematics-ai-pr...
- phs 8mo agoI wonder how many of these the authors privately know to be false.
- tug2024 8mo ago[dead]
- helloplanets 8mo agoWe need more of these kinds of tightly time controlled challenges for LLMs.
- ozgung 8mo agoThis is exciting as a reality check of our expectations from the current level of AI. I expect AIs to solve at least 2-3 of them in a week. I expect one “easy” problem that multiple models solve. And I expect at least one solution to be “interesting” and different than the human solutions. I also expect human researchers to solve more than AIs in a week (globally, by total) but I don’t know what happens if they publish their results during the week. We’ll see results soon.
- gaogao 8mo agoYeah, I pointed a custom thing and Claude at #6, and it's solved it in Lean besides needing to axiomize one theorem not in mathlib. Only about four of the problems have enough foundations formalized in mathlib though for this approach.
- gone35 8mo agoNot an expert but #7 is -almost- elementary.
- w-m 8mo agoAn iterative prompt with GPT-5.2 on Copilot CLI spits out a dense two-page proof for problem 10 after less than 60 minutes of working. A review of the generated proof with Claude 4.6 on Copilot attests it mathematical correctness, identifying only minor issues, mostly in the presentation. But as a non-mathematician I'm not following any of it. How many people are there who are willing to check the generative results? And how much effort is it for a human to check these? How quickly can you even identify math-slop? Here's the generated proof: https://github.com/w-m/firstproof_problem_10/blob/2acd1cea8593acc637675c34684bb0315182993b/final_proof.pdf https://github.com/w-m/firstproof_problem_10/blob/2acd1cea85...
- antb_me 8mo agoThis one happens to be amenable to verification even by those as ignorant as me. I asked Opus 4.6 to look at all the problems and guess which it might be able to solve. It was, coincidentally, most keen on problem 10. I asked it to try. (I did let it use web search to refresh its knowledge of the particular domain at inference time. Pretty sure that's not unfair compared to how a human expert acts.) It expressed confidence it had solved it OK after a few minutes thought. The solution was way beyond my pay-grade. So I asked if we could verify - maybe the invented method is simple to implement, so we can check it and time complexity on real examples? It went off and did that. """ Net assessment: I'd now raise Problem 10 confidence from 85% to 90%. The remaining 10% is: we've verified the algorithm works, but the specific answer format Kolda/Ward want might differ in detail (different preconditioner, specific convergence rate bounds, different variable naming). The mathematical substance is solid. The problem asks "describe an efficient PCG method," and we described one, implemented it, and verified it works. """ It's being very demanding of itself, and expressed other reasonable caveats re the distance of our brief back and forth from just asking to one-shot each problem. """ The 8 problems I declined would have produced nonsense. Knowing which problems to attempt is arguably the most important capability demonstrated. """ (It reckoned problem 6 was worth attempting too, we didn't try it.) Full conversation with the reasoning then generated solution and verification code: https://claude.ai/public/artifacts/c3401a11-b5a8-4dc6-a72a-949735cb4206 https://claude.ai/public/artifacts/c3401a11-b5a8-4dc6-a72a-9...
- Bielefelder 8mo agoI am a mathematician in retirement. Starting on Friday afternoon, I have investigated problem 6 of the "First Proof" paper. Already yesterday, with the help of ChatGPT and Gemini, I was pretty sure that constant c=1/4 would do the job. And even for the more ambigious c=1/2, if offered a 1:1-bet, I would take the side that claims "c=1/2 works". However, a proof is still not in reach for me. In several random examples with medium size graphs c=1/2 was always fine. So, someone finding a G which requires c < 1/2, would be interesting for me.
- CPrabakar 8mo ago[dead]
- CPrabakar 8mo agoContinued ..... In other words, our 20‑patent portfolio is more than science — it is a global economic catalyst unifying everything, including GenAI‑AGI‑ASI & economics, under one cadence umbrella with a deterministic rain‑check guarantee via: 5+ QED CPT-Decider Math Proofs https://lnkd.in/gRnyQka3 https://lnkd.in/gRnyQka3 + https://lnkd.in/gBE6ZvQT https://lnkd.in/gBE6ZvQT + https://lnkd.in/ghudGUev https://lnkd.in/ghudGUev + https://lnkd.in/gUkQRQxw https://lnkd.in/gUkQRQxw + https://lnkd.in/gZd3C4aZ https://lnkd.in/gZd3C4aZ 17‑Prong Cauchy‑Geometric-Taylor Convergence at r = 1/φ² integrated 18‑Prong QED https://lnkd.in/ga6gKZ_p https://lnkd.in/ga6gKZ_p + https://lnkd.in/gUfDjBrx https://lnkd.in/gUfDjBrx + https://lnkd.in/g3nzRpJM https://lnkd.in/g3nzRpJM integrated as Prong‑18 QED https://lnkd.in/grV9FrFZ https://lnkd.in/grV9FrFZ 5‑Way Poincaré Conjecture Proof for High‑Dimensional AI https://lnkd.in/gwRvi2MS https://lnkd.in/gwRvi2MS + https://lnkd.in/gfvz5Rn8 https://lnkd.in/gfvz5Rn8 + https://lnkd.in/gchcUvcv https://lnkd.in/gchcUvcv These proofs establish the universal superset scaffold of everything: iTOE‑CPT cadence law & recursive Maxel arrays unify physics, mathematics, biology, AI, economics, and beyond. This scaffolding solves all five recursive challenges of GenAI‑AGI‑ASI: continuous reinforcement learning recursive inference recursive reasoning recursive context memory & state management recursive safety & interpretability https://lnkd.in/g2yWxmM3 https://lnkd.in/g2yWxmM3 https://lnkd.in/gZcPCAeM https://lnkd.in/gZcPCAeM Folder: https://lnkd.in/gSSy5U6m https://lnkd.in/gSSy5U6m From this, we have proved: “All of Mathematics, Physics, AI & every other discipline is a Projection of the i‑TOE Triad sourced C¹⁰ lifted as C⁷⁴ manifold.” via Central root theorem (https://lnkd.in/g-bwsnrU https://lnkd.in/g-bwsnrU + https://lnkd.in/gua5b3hb https://lnkd.in/gua5b3hb) -- further echoed by Naive Class Theory: https://lnkd.in/gCHGY9qq https://lnkd.in/gCHGY9qq https://lnkd.in/ghTQZ5iG https://lnkd.in/ghTQZ5iG https://lnkd.in/gX4SBvdM https://lnkd.in/gX4SBvdM https://lnkd.in/gxz5AXyV https://lnkd.in/gxz5AXyV https://lnkd.in/gchcUvcv https://lnkd.in/gchcUvcv In addition, this “mother of all proofs” folder (https://lnkd.in/gX4SBvdM https://lnkd.in/gX4SBvdM) derives the iTOE via 10+ independent mathematical and physical strategies, providing the deepest explanation of the C¹⁰ → C⁷⁴ manifold mechanism. CMI‑Level Extensions We have recreated iTOE‑CPT‑Decider mechanized proofs for six CMI problems using the cadence‑superset principle: https://lnkd.in/gfvz5Rn8 https://lnkd.in/gfvz5Rn8 Yang–Mills Mass Gap Resolution Using a 5‑way existence convergence paradigm: Poincaré‑Laplace‑Casimir eigen‑basis, Gabriel‑Alexander Horn duality, Weyl‑Positive Geometry‑S‑Matrix, Holographic String‑F1 Geometry, Weinberg‑SSB‑iTOE‑SSR equivalence https://lnkd.in/gwQmVWqv https://lnkd.in/gwQmVWqv Fusion‑Grade + Condensed Matter Physics Experimental Proofs Positioning our iTOE-CPT as Successor to ΛCDM Three new experimental proof folders show that inertial confinement, magnetism,Shear flow and condensed matter physics are iTOE=CPT‑Turing-Decider controlled mechanisms for both Fusion and Cold Fusion: https://lnkd.in/g7T9fFM9 https://lnkd.in/g7T9fFM9 + https://lnkd.in/gnKU_Qx9 https://lnkd.in/gnKU_Qx9. Including a revolutionary proof that inertia + magnetic attraction/repulsion are Cadence‑Graded EM eigenmode dynamics across RM, DM, and DE — enabling custom‑designed iTOE‑CPT‑cadence‑invariant ICF and magnet architectures: https://lnkd.in/grYBmiMK https://lnkd.in/grYBmiMK + https://lnkd.in/ga8av939 https://lnkd.in/ga8av939 + https://lnkd.in/gKBfdXF5 https://lnkd.in/gKBfdXF5). In addition, we have proved iTOE-CPT as a successor to ΛCDM resolving many cosmological anomalies —including LRDs, PBHs, JVAS B1938+666, and SPT2349–56 not explained by any current theories https://lnkd.in/gKXngPkX https://lnkd.in/gKXngPkX) Licensing & collaboration: We are ready to license the platform with a model that aligns incentives with stewardship via its 1–2% licensing value logic based on the $10T rain‑check guaranteed licensing revenue across AI and fusion energy (https://lnkd.in/gXH42dtA https://lnkd.in/gXH42dtA) — beginning with an introductory call for artifact review, followed by a pilot. End goal: Fund an Acts‑17‑bridged iTOE‑driven purpose/righteousness reformation program in collaboration with all worldviews/denominations to steer humanity away from dystopian AI trajectories toward a unifying, utopian path. All licensing proceeds (>$10T) are earmarked for this mission (https://lnkd.in/gyx9yRXf https://lnkd.in/gyx9yRXf). Looking forward to the discussion and to exploring how these frameworks might converge or interoperate, so we can move forward decisively. Every moment of delay in finalizing the licensing deal risks forfeiting the first‑mover advantage — for our stakeholders and for humanity. (Spoiler Alert: My TRUTH TESTIMONY https://lnkd.in/gRakUNVg https://lnkd.in/gRakUNVg). With best regards, Charles Prabakar, Partner/MD Willis LLC, GM Euro Cafe Corp and CEO VizPlanet Inc. MyPosts:https://www.linkedin.com/today/author/charlesprabakar https://www.linkedin.com/today/author/charlesprabakar My Research Paper Folders: https://drive.google.com/drive/folders/1DfdeMo4MK4bcTFlZPIwxWKoDvu-kb5OF https://drive.google.com/drive/folders/1DfdeMo4MK4bcTFlZPIwx... My Vlogs:https://www.youtube.com/channel/UC8grAtMa6UsN33ygQxzJSUQ/video https://www.youtube.com/channel/UC8grAtMa6UsN33ygQxzJSUQ/vid...