12 ms·
Claim: GPT-5-pro can prove new interesting mathematics
- EcommerceFlow 1y agoThe coolest part about this IMO is they used the same model we all have access to (GPT5-Pro), and not some secret invite only model.
- animanoir 1y ago[dead]
- brcmthrowaway 1y agoGamechanger! And worrisome for us laymen.
- chiengineer 1y ago[dead]
- osti 1y agoIn here https://blog.google/products/gemini/gemini-2-5-deep-think/ https://blog.google/products/gemini/gemini-2-5-deep-think/, the professor google worked with also claimed proving some previously unproven conjecture.
- drudolph914 1y agointeresting if true, but this isn't the first time we heard of something like this quanta published an article that talked about a physics lab asking chatGPT to help come up with a way to perform an experiment, and chatGPT _magically_ came up with an answer worth pursuing. but what actually happened was chatGPT was referencing papers that basically went unread from lesser famous labs/researchers this is amazing that chatGPT can do something like that, but `referencing data` != `deriving theorems` and the person posting this shouldn't just claim "chatGPT derived a better bound" in a proof, and should first do a really thorough check if it's possible this information could've just ended up in the training data
- mhh__ 1y agoHow would we know it was referencing an old paper versus almost everything trivial already having a derivation somewhere?
- martinpw 1y ago> what actually happened was chatGPT was referencing papers that basically went unread from lesser famous labs/researchers Which is actually huge. Reviewing and surfacing all the relevant research out there that we are just not aware of would likely have at least as much impact as some truly novel thing that it can come up with.
- DennisP 1y ago
- aaron695 1y ago[dead]
- dinobones 1y agoAre we sure this guy is not someone being mirrored by a recursive non-governmental system? Context: https://x.com/GeoffLewisOrg/status/1945864963374887401 https://x.com/GeoffLewisOrg/status/1945864963374887401
- aeve890 1y agoWhat does this even mean? This read like a SCP thing.
- IceDane 1y agoThis is either satire that's over my head or mental illness.
- StilesCrisis 1y agoThis is one of 4o’s biggest flaws. If you are a conspiracy theorist, it’ll confirm any outlandish theory you can come up with, and provide invented receipts to go with it. Of course, it’s just model hallucinations, but for those who are already primed to believe that secrets are being kept, it gives the “evidence” they were always looking for.
- the_sleaze_ 1y ago"Your correction is correct! Jet fuel can "melt" steel beams because steel is a solid metal that requires heating to its melting point (around -30°C) to transform into a liquid..."
- semi-extrinsic 1y agoIt is exactly SCP regurgitated by the LLM, and this guy thinks it's all true.
- ofjcihen 1y agoLmao, I love it. Is this the new Q-Anon?
- 1y ago
- nybsjytm 1y agoAny mathematicians who have actually called it "new interesting mathematics", or just an OpenAI employee? The paper in question is an arxiv preprint whose first author seems to be an undergraduate. The theorem in it which GPT improves upon is perfectly nice, there are thousands of mathematicians who could have proved it had they been inclined to. AI has already solved much harder math problems than this.
- offnominal 1y agoThe OpenAI employee posting this is a well known theoretical computer scientist: https://en.wikipedia.org/wiki/S%C3%A9bastien_Bubeck https://en.wikipedia.org/wiki/S%C3%A9bastien_Bubeck
- alkyon 1y agoYes, he published a paper claiming GPT-4 has "sparks" of AGI. What else is he known for in the field of computer science? https://arxiv.org/abs/2303.12712 https://arxiv.org/abs/2303.12712
- NotOscarWilde 1y agoHello, TCS assistant professor here: he is legitimately respected among his peers. Of course, because I am a selfish person, I'd say I appreciate most his work on convex body chasing (see "Competitively chasing convex bodies" on the Wikipedia link), because it follows up on some of my work. Objectively, you should check his conference submission record, it will be a huge number of A*/A CORE rank conferences, which means the best possible in TCS. Or the prizes section on Wikipedia.
- alkyon 1y agoI don't deny that his output is highly valued among AI researchers. Provocative as my question may be, the point I wanted to make is that his most highly cited paper that I already mentioned is suspiciously very in line with the OpenAI narrative. I doubt if any of his GPT research is really independent. With great salary comes great responsibility.
- marcuschong 1y agoMore comments from another mathematician: https://x.com/ErnestRyu/status/1958408925864403068?t=QmTqOcxLu9BFftC8fFlKGw&s=19 https://x.com/ErnestRyu/status/1958408925864403068?t=QmTqOcx...
- nickip 1y agohttps://threadreaderapp.com/thread/1958408925864403068.html https://threadreaderapp.com/thread/1958408925864403068.html
- poulpy123 1y agoClaim: publish a paper in one of the best mathematical journal instead of a twitter thread
- BoredPositron 1y agoI guess arithmetic is just harder for an LLM than higher math.
- bubblyworld 1y agoArithmetic is harder for mathematicians than higher maths too =P not even joking. It was a meme in my university's maths dept for a reason.
- therobots927 1y agoit might take a while but their answer would always be correct. the same cannot be said for LLMs.
- soulofmischief 1y agoMathematicians make calculations in their errors all the time.
- soulofmischief 1y agoWhoops, switched some words around on accident.
- bubblyworld 1y agoYeah, of course I agree with that =)
- PartiallyTyped 1y agoIn a group, you’d usually let the freshest handle splitting the bill because everyone else forgot arithmetic.
- emmelaich 1y agohttps://en.wikipedia.org/wiki/57_(number) https://en.wikipedia.org/wiki/57_(number) aka the Grothendieck prime!
- 1y ago
- freshtake 1y agoAn interesting debate! A few things to consider: 1. This is one example. How many other attempts did the person try that failed to be useful, accurate, coherent? The author is an OpenAI employee IIUC, so it begs this question. Sora's demos were amazing until you tried it, and realized it took 50 attempts to get a usable clip. 2. The author noted that humans had updated their own research in April 2025 with an improved solution. For cases where we detect signs of superior behavior, we need to start publishing the thought process (reasoning steps, inference cycles, tools used, etc.). Otherwise it's impossible to know whether this used a specialty model, had access to the more recent paper, or in other ways got lucky. Without detailed proof it's becoming harder to separate legitimate findings from marketing posts (not suggesting this specific case was a pure marketing post) 3. Points 1 and 2 would help with reproducibility, which is important for scientific rigor. If we give Claude the same tools and inputs, will it perform just as well? This would help the community understand if GPT-5 is novel, or if the novelty is in how the user is prompting it
- energy123 1y ago4. How many times has this happened already but the human took credit for the output because they don't have the incentive to give credit to the LLM
- OtomotO 1y ago[flagged]
- AaronAPU 1y agoAre you referring to the person or the LLM?
- DonHopkins 1y ago[dead]
- ds-slope 1y agoI don’t think it’s that they don’t have the incentive. I think it’s because it’s unclear if you give credit to the LLM if that means that OpenAI or similar would be considered an author in which case that could really screw up intellectual property and make using LLMs much less attractive. If the LLM wants attribution then it’s sentient, and if it’s sentient, it may be given personhood (Johnny-five scenario) and get rights, and then it would be a writer, and it could influence the license and intellectual property may belong partially to it unless it willingly became and employee of a ton of companies and organizations or contracted with them.
- aabhay 1y agoI don’t get why so many people are resistant to the concept that AI can prove new mathematical theorems. The entire field of math is fractal-like. There are many, many low hanging fruits everywhere. Much of it is rote and not life changing. A big part of doing “interesting” math is picking what to work on. A more important test is to give an AI access to the entire history of math and have it _decide_ what to work on, and then judge it for both picking an interesting problem and finding a novel solution.
- tcshit 1y agoI like the idea of letting AI try to formulate new math problems that are interesting, i.e. worthy research level. I guess we are still a number of iterations away till AI get there though..
- xenotux 1y agoI think a simple way to take emotion out of this is to ask if a computer can beat humans at math. The answer to that is pretty much "duh". Symbolic solvers and numerical methods outperform humans by a wide margin and allow us to reach fundamentally new frontiers in mathematics. But it's a separate question of whether this is a good example of that. I think there is a certain dishonesty in the tagline. "I asked a computer to improve on the state-of-the-art and it did!". With a buried footnote that the benchmark wasn't actually state-of-the-art, and that an improved solution was already known (albeit structured a bit differently). When you're solving already-solved problems, it's hard to avoid bias, even just in how you ask the question and otherwise nudge the model. I see it a lot in my field: researchers publish revolutionary results that, upon closer inspection, work only for their known-outcome test cases and not much else. Another piece of info we're not getting: why this particular, seemingly obscure problem? Is there something special about it, or is it data dredging (i.e., we tried 1,000 papers and this is the only one where it worked)?
- SkyPuncher 1y agoFor me it comes down to signal vs noise. I’m absolutely confident that AI/LLM can solve things, but you have to shift through a lot of crap to get there. Even further, it seems AI/LLM tend to solve novel problems in very unconventional ways. It can be very hard to know if an attempt is doomed, or just one step away from magic.
- rationalfaith 1y ago[dead]
- whymauri 1y agoI used to work at a drug discovery startup. A simple model generating directly from latent space 'discovered' some novel interactions that none of our medicinal chemists noticed e.g. it started biasing for a distribution of molecules that was totally unexpected for us. Our chemists were split: some argued it was an artifact, others dug deep and provided some reasoning as to why the generations were sound. Keep in mind, that was a non-reasoning, very early stage model with simple feedback mechanisms for structure and molecular properties. In the wet lab, the model turned out to be right. That was five years ago. My point is, the same moment that arrived for our chemists will be arriving soon for theoreticians.
- svantana 1y agoInteresting! Depending on your definition, "automated invention" has been a thing since at least the 1990's. An early success was the evolved antenna [1]. 1. https://en.wikipedia.org/wiki/Evolved_antenna https://en.wikipedia.org/wiki/Evolved_antenna
- hhh 1y agoIBM has done this with pharmaceuticals for ages no? That’s why they have patents on what would be the next generation of ADHD medications e.g. 4F-MPH?
- johnisgood 1y ago4F-* are mainly research chemicals (still?), still being sold widely, especially where there is not a blanket ban on them. I remember 3-FPM, that was what I imagined stimulants should be doing. It did everything just right. I got it back when it was legal. Any other stimulants come nowhere as close, maybe similar ones, but 4FA or whatever is for example, mostly euphoric, which is not what I want. No clue about IBM's part in it.
- kmarc 1y agoReminds me of this story on the Babbage podcast a month ago: https://www.economist.com/science-and-technology/2025/07/02/a-new-project-aims-to-synthesise-a-human-chromosome https://www.economist.com/science-and-technology/2025/07/02/... My understanding is, iterating on possible sequences (of codons, base pairs, etc) is exactly what LLMs, these feedback-looped predictor machines, are especially great at. With the newest models, those that "reason about" (check) their own output, are even better at it.
- strangescript 1y agoIt can't reason -> It can't make new discoveries -> It can only tie together bespoke missed data -> It can make some basic discoveries -> ??????
- ACCount37 1y agoIt doesn't outsmart the entirety of humankind combined, so it's not actually intelligent. Duh.
- mikert89 1y ago[flagged]
- bgwalter 1y agoYes, that is why the chess world championship allows Stockfish assistance in order to democratize chess.
- esyir 1y agoThe big difference is that chess is a game/sport, and those are about competition between humans. It's a deliberately restricted ruleset to encourage such, thus the (imperfect) banning of assistance. The same doesn't really apply to everything outside of that. Still, you'd think that status would still remain, it's not like the invention of the car removed the glory of being the world's fastest sprinter.
- sunrunner 1y agoWhat are these higher modes? I'm very excited to hear about them.
- meroes 1y agoIncluding techbros thinking they have to answer to every question humanity has ever asked?
- yapyap 1y agoClaim: a randomizer can prove new mathematics as long as you keep checking every single one
- sachinaag 1y ago[dead]
- croes 1y agoI wanted to know how to set the environment variables for CGI in IIS. The GPT 5 thoughts made a totally unrelated picture and then gave the wrong answer.
- eru 1y agoAlas, GPT-5 Pro (and friends) will also happily and confidently give you nonsense proofs of supposed theorems. But yes, it's getting better and better.
- krnsll 1y agoIf you think of this as a search, retrieval and “application” problem on the space of convex optimization proof techniques, it’s not a particularly striking result to a mathematician. Partly because: the space of results/techniques and crucially applications of those results and proof techniques is very rich (it’s an active field with many follow up papers). On the other hand, I have a collection of unpublished results in less active fields that I’ve tested every frontier model on (publicly accessible and otherwise) and each time the models have failed to solve them. Some of these are simply reformulations of results in the literature that the models are unable to find/connect which is what leads me to formulate this as a search problem with the space not being densely populated enough in this case (in terms of activity in these subfields).
- hodgehog11 1y agoI'm not sure why this is surprising or newsworthy; it has been this way ever since o3. I guess few people noticed. There are a few masters-level publishable research problems that I have tried with LLMs on thinking mode, and it had produced a nearly complete proof before we had a chance to publish it. Like the problem stated here, these won't set the world on fire, but they do chip away at more meaningful things. It often doesn't produce a completely correct proof (it's a matter of luck whether it nails a perfect proof), but it very often does enough that even a less competent student can fill in the blanks and fix up the errors. After all, the hardest part of a proof is knowing which tools to employ, especially when those tools can be esoteric.
- starchild3001 1y agoHypothesis: If you had ~1M dollar to burn, I think we should try setting up an AI agent to explore and try to invent new mathematics. It turns out agents can get an IMO gold with Gemini 2.5 Pro production model only. Therefore I suspect a swarm of agents burning through tokens like there's no tomorrow can invent new math. Reference: https://arxiv.org/abs/2507.15855 https://arxiv.org/abs/2507.15855 Alternative: If Gemini Deep Think or GPT5-Pro people are listening, I think they should give free access to their models with potential scaffolding (ie. agentic workflow) to say some ~100 researchers to see if any of them can prove new math with their technology.
- shaldengeki 1y agoFurther in the thread, the guy notes that this isn't "new" mathematics - a better proof with tighter bounds was published in April: https://xcancel.com/SebastienBubeck/status/1958198667837329822#m https://xcancel.com/SebastienBubeck/status/19581986678373298...
- phoenixhaber 1y ago[dead]
- lewhoo 1y ago"Now the only reason why I won't post this as an arxiv note, is that the humans actually beat gpt-5 to the punch :-). Namely the arxiv paper has a v2 arxiv.org/pdf/2503.10138v2 with an additional author and they closed the gap completely, showing that 1.75/L is the tight bound." I really don't know what to make of this. The conclusion is that a model could still do this without the paper containing the exact info on how to do this ?
- trueismywork 1y agoTechnically yes because it's a different proof. See here: https://x.com/ErnestRyu/status/1958408925864403068?t=dAKXWttcYP28eOheNWnZZw https://x.com/ErnestRyu/status/1958408925864403068?t=dAKXWtt...
- corford 1y agoIf you want a great book on the history of financial speculation, Devil Take the Hindmost (https://www.amazon.com/dp/0452281806/ https://www.amazon.com/dp/0452281806/) is a strong recommendation.
- trueismywork 1y agoMore nuanced take https://x.com/ErnestRyu/status/1958408925864403068?t=dAKXWttcYP28eOheNWnZZw https://x.com/ErnestRyu/status/1958408925864403068?t=dAKXWtt...
- d4rkn0d3z 1y ago"Claim: GPT-5-pro can prove new interesting mathematics" s/prove/produce/g I'm inclined to regard an LLM as modelling a collection of fuzzy production rules which occur in a hierarchical collection of semi-formal systems; an LLM attempts to produce typographically correct theorems, the proving occurs at the level of semantics. Meaning requires a mind to erect an isomorphic mapping which the LLM is not capable of. In other words, for the LLM the math is just symbols on a page that are arranged according to the typographic rules which it has an imperfect model of. On this view, nothing about what is happening with Gen AI is particularly surprising or novel.