16 ms·
Ten advances in mathematics and theoretical computer science
- joshlk 2mo agoSome of the Lean proofs are 50k lines - is that normal?
- piker 2mo agoI don’t feel the existential dread of mathematicians is correct. It seems to me in fact these results are bringing math mainstream. I now personally look forward to the interpretations and discussions of the significance of such results by human mathematicians. Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians. [edit: deleted a distracting comparison to Chess]
- aabhay 2mo agoGiven that we were nowhere near this state even two years ago, I think it’s a question of velocity more so than just distance.
- traes 2mo agoEvery time someone makes a comparison to chess I die inside. Chess is a spectator sport primarily funded by a few eccentric billionaires. Players artificially constrain themselves in timed environments knowing that they will never be able to produce better moves than a smartphone because a select few people find it interesting. Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs. I shudder to imagine what will happen to the tens of thousands of non-Fields medalist caliber mathematicians if math goes the way of chess. Perhaps Terence Tao and a few other famous mathematicians will be funded by Peter Thiel to report on how well humanity can keep up with the machines? How do you expect any mathematician to be optimistic about this comparison.
- anematode 2mo agoFully agreed. As someone who both loves chess and works on chess engines... these comparisons to chess needs to stop.
- ratmice 2mo agoAnother noteworthy difference is that Stockfish is also gpl.
- traes 2mo agoIf there was any real money in it Stockfish would not be the best chess engine.
- ratmice 2mo agoThats not the point, if there were a better proprietary engine stockfish would still be there as a baseline. Anyone can access an engine as good as stockfish to practice against. Are any open models touting mathematical breakthroughs?
- traes 2mo agoThere is money in this, so of course the closed models are far ahead. The open models will likely catch up a bit at some point, just as Stockfish caught up to AlphaZero. That being said, there are already a couple. It seems Deepseek has a claimed proof to the "Ziegler's Cross-Polytope Conjecture" [0], but I can't speak to the significance of the result. [0] https://arxiv.org/abs/2606.31640 https://arxiv.org/abs/2606.31640
- energy123 2mo agoThe distinction is mathematician vs mathematics. Mathematics is going to reach new heights beyond the wildest dreams of contemporary mathematicians. But perhaps without the participation of many paid mathematicians.
- 2mo ago
- baq 2mo agoAs in chess and go and also coding for the past ~year there are two groups of people: the disappointed and the enthusiastic. The disappointed are sad that they lost their advantage and that the craft they honed for years or decades has rapidly lost its value; the enthusiastic are excited about the future and what computers can bring to their domain and how it will evolve. I’m a bit of both if it comes to programming, more enthusiastic than disappointed, but also more than a bit terrified about the pace of it all. I imagine that’s how Kasparov felt back then, that’s how Lee Sedol felt and now that’s how Terry Tao feels. The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.
- traes 2mo agoA fundamental difference being that no one was actually paid to find good moves in chess and go like they are to solve math problems and write code. You're comparing the digital camera and the automobile.
- jibal 2mo agoThe chess analogy is awful. If you simply want to know the answer to a chess problem, give it to the engine. Chess only lives on because it's a competition between humans to test their skill (just like bicycles, cars, trains didn't eliminate foot races) ... the computer is largely factored out, but not entirely -- people train with the computer, use it to check whether they played correctly, ... and they cheat. A lot. Thus there are more and more sophisticated mechanisms to detect and prevent cheating. If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof). P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.
- piker 2mo agoI’ve deleted it but no it’s not awful anymore than saying “we survived WWII, we can survive this.” The point was that change happens but humans find a way forward.
- energy123 2mo agoThe old way of establishing career credibility is being destroyed, for better or worse. Accomplishments that used to be career-defining are hard to distinguish from AI, and correlate more with access to compute. Think about Bill Gates's math paper he wrote in college. That kind of thing is gone now as a path to credibility. There's still competitions and grades, but the diversity of paths is going away. Maybe new ones will open up. This is a competitive advantage for old people who have credible pre-2025 accomplishments they can point to.
- traes 2mo agoIf accomplishments can't be distinguished between talented people and untalented people with compute, is there really a point in trying? I suppose one can hope that talented people given compute will be more effective than untalented people with compute, but I despair that that may not be true for much longer.
- deleted 2mo ago[deleted]
- FranzFerdiNaN 2mo agoKnowledgable people can confirm what the AI produces is correct. I could make ChatGPT produce a result on an open question and I would have zero way to verify its actual correctness. Which is less interesting work. And you probably need to do the hard grunt work by hand first to develop the skills and intuition to be able to verify an AI-generated result. So you can’t outsource everything to AI without loss of skill.
- deleted 2mo ago[deleted]
- noslenwerdna 2mo agoI mean there were problems with the "old way" as well. Not clear if this change is net positive or negative in my opinion.
- kzrdude 2mo agoDo mathematicians have the right to say "no AI PRs please, the volume is too much" just like how some open source maintainers do it? I guess they feel a loss of control, there is no way to turn the hose off. Thinking of this a little bit with the perspective of every new proof as a burden, dumped for review by actual mathematicians.
- dash2 2mo agoI find this whole way of looking at things weird. Did maths exist just to entertain and employ mathematicians? Surely maths is, like, useful? Not immediately, not predictably, but in the long run? In which case, whether mathematicians feel bad about it is mostly irrelevant - it's like complaining about the railway because it may put coaching inns out of business.
- MinimalAction 2mo agoAbsolutely not the same! People need jobs to bring in income. I don't believe those who profit off of this will share it with the world. The power is all concentrated in the few hands that decide whether or not the rest get any semblance of income in the long run. I don't believe UBS until it happens.
- dash2 2mo agoIt sounds like you think no technological advances will make the world richer in the long run. I politely suggest that the past century of economic growth shows problems with this argument. I also think that, while jobs are important for prosperity, the jobs of mathematicians are a minuscule fraction of a percent of the total.
- MinimalAction 2mo agoSo, you're arguing that just because mathematicians are a minuscule, it doesn't matter for prosperity? Because, if so, that is an insane take honestly.
- dash2 2mo agoReally? "The benefit to all of humanity from advances in mathematics outweighs the lost jobs of research mathematicians". That's an insane take?
- MinimalAction 2mo ago
- c0rruptbytes 2mo agoi think i agree, we are going to have /more/ math and now need /more/ mathematicians (we are seeing https://vibemathed.com/ https://vibemathed.com/) these LLMs are great are generating arguments but they don't ask questions, we will need mathematicians to shepherd them into more discoveries i really want to see open weight models crack some breakthroughs
- andai 2mo ago> LLMs don't ask questions Why don't they? That sounds like an important problem to solve. Along with the fact that they can't learn anything (after the training stops).
- danielrmay 2mo agoI'm enjoying learning about these hard problems, but this line about credit made me chuckle: > We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
- emil-lp 2mo agoNo, the correctness isn't for the "inside the Lean proofs", but for the translation of "human language math" and its formal Lean variant.
- danielrmay 2mo agoI see. It still feels like a bit of an oddly solemn way of saying "this is the part we admit responsibility for"
- emil-lp 2mo agoWell, to be fair, with Lean proofs, that's the only thing there is (unless I'm missing something).
- baq 2mo agoIt’s more than you get from free software - you get no proofs, no warranties and any responsibility of its authors are their pure good will. Reminder lean proofs are software!
- jhanschoo 2mo agoTraditionally, a mathematician would be implicitly responsible for all that (if they were to publish Lean code) and also the intellectual work that led to the artifact of the mathematical paper (and code, if part of the contribution). This statement should rather be read as an acknowledgement of limitation of authorship from the implicit, traditional understanding.
- traes 2mo ago
- emil-lp 2mo agoI wonder what the total cost of this research was, including the salary for their mathematicians and engineers.
- z7 2mo ago> The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices. https://x.com/polynoamial/status/2083470822258467194 https://x.com/polynoamial/status/2083470822258467194
- traes 2mo agoGiven that OpenAI pays their employees with stock surely a breathtaking number, but not a very meaningful number now that the infrastructure is in place and the models are trained. AI could never get better and it would still be incredibly disruptive.
- kingstnap 2mo agoWhy would you factor in salary unless they had to baby it through. You would only count the hours for setting up the harness and prompt and checking the result. Training the model is going to be amortized over other uses.
- emil-lp 2mo ago> Why would you factor in salary Say that it turned out that the total cost of the proof of the Erdős unit-distance conjecture was $50 million. Then the question really becomes: yes, these models are capable of proving important mathematical results, but at a very high cost. Is it worth it? If a mathematician applied for a research grant of $50M USD for proving the same thing, they would have been laughed out of the bank. What's more is that when you have a research grant, you train PhDs and postdocs, you hire new staff, and you disseminate. That is, you get much more value for the money spent. I'm just curious what the cost is.
- aabhay 2mo agoMy main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup. I want to know: 1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
- einpoklum 2mo agoAlso, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar? Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
- traes 2mo ago> Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar? A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits. > Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work. I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof. [0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais_cdc_proof_announcement_gpt56_used_a/ https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...
- irthomasthomas 2mo agoWhy you think that?
- 0x5FC3 2mo agoHow much do you all think it would cost to "buy" these advances from PhDs, practicing scientists?
- traes 2mo agoThis isn't really a productive way to think about these things, IMO. It's quite possible it would take hundreds of years for any specific group of PhDs to solve them. Or one individual PhD could have the correct flash of insight and solve it in a month. There's absolutely no way to predict this, besides trying to gauge the apparent simplicity of the proof or counterexample (which is likely to be misleading). Until someone actually runs an experiment like this it's not a viable metric.
- 0x5FC3 2mo agoI understand and I am not trying to deny the impressiveness or the velocity of AI in general. But at some point we have to ask how much do we trust the labs at face value without much transparency of how they got to the results when there is trillions of dollars on the line.
- simianwords 2mo agoThe level of conspiracy theory is nuts
- 0x5FC3 2mo agoI would say the lack of skepticism is nuts, honestly.
- frozenseven 2mo agoCapabilities of this sort have already been demonstrated by independent parties, and models have consistently gotten better at this. Yes, insinuating that mathematicians and scientists are secretly solving decades-old problems on OpenAI's behalf is an insane conspiracy theory.
- zkmon 2mo ago> claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work. AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way. Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
- NitpickLawyer 2mo agoA better analogy would be a manufactured object, say 3d printed for simplicity. The 3d printer is given an input, and an object manifests itself after some time. We say that the creator of the object is the person turning on the machine, sending the data, and collecting the object. Not the machine itself.
- cure_42 2mo agoI'd say the creator is the one who created the 3d model, not the one who pushed the print button.
- NitpickLawyer 2mo ago(let's assume that)My 3dprinter is special. It has a bunch of values + an algorithm (i.e. a neural network) that takes input as tokens and outputs a printed object.
- dgellow 2mo agoI would say „I made this gadget with my 3d printer, but the designer is someone else (I found the model online)“. The intent, the drive, the action comes from the human
- 2mo ago
- deleted 2mo ago[deleted]
- s_Hogg 2mo agoI don't know why, but when I saw the source of this particular headline it reminded me of the album title 26 Mixes for Cash
- defrost 2mo agoAmbient 0: Math for Airports
- utopiah 2mo ago[flagged]
- utopiah 2mo agoTo clarify a bit due to the downvotes : this is not a research paper from a startup or a public frontier lab, it is just PR from a corporation, thus yes an advertisement. Downvote all you like it's still of no value.
- Windchaser 2mo ago> this is not a research paper from a startup or a public frontier lab, it is just PR from a corporation If they're publishing the solutions to these 10 problems, then this is, essentially, the announcement of 10 research papers. Yes, it's partly for reputation (as are many research papers), but that doesn't mean it's of no value. The best way to advertise is to show that you're providing value.
- luciana1u 2mo ago[flagged]
- baq 2mo agoI asked ChatGPT and it told me these aren’t not important /s
- avaer 2mo agoWhat happens when OpenAI et al stop being open about these things, and just pack it into the training?
- traes 2mo agoNot much point to pure math being kept secret, in all honesty. There isn't really industrial value, its only purpose (to them) is showing off their model's capabilities. More realistically they'll just stop paying for it. Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.
- asdewqqwer 2mo agoAt this stage. No doubt calculus had plenty industrial benefit.
- deleted 2mo ago[deleted]
- simianwords 2mo agoWhat does this even mean lol. These are not solved questions. The solution never existed.
- deleted 2mo ago[deleted]
- sillysaurusx 2mo agoHey man, just wanted to say hi and catch up a bit. I tried DMing you on Twitter. If that sounds interesting then shoot me a message sometime. Hope you’ve been well :)
- sergiomiguens 2mo agohttps://mathstodon.xyz/@sergiosh/116847266951670037 https://mathstodon.xyz/@sergiosh/116847266951670037 https://mathstodon.xyz/@sergiosh/116976232522229829 https://mathstodon.xyz/@sergiosh/116976232522229829 Look at this two threads.
- lifeisstillgood 2mo agoOn the token limits etc - one assumes that OpenAI et al are able to “hire expert in field, and let them spend the equivalent of a million dollars of tokens” because they are not actually selling their complete compute 24 hrs a day, so the cost internally is a negligible (ish) electricity bill. Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
- traes 2mo agoPresumably it's a rounding error compared to their full output, and they're making sure they have enough compute set aside for research by limiting public models. The more datacenters they build the less they have to limit them.
- lwansbrough 2mo agoFor OpenAI, research is marketing. I’m sure they’ve got plenty of budget for that.
- Davidzheng 2mo agoRL training can use all of them - idk what needed means.
- simianwords 2mo agoI love how people come up with creative ideas to prove the bubble. This one is even more ridiculous - that OpenAI had spare compute to advance mathematics proves that data centres will not be needed. WHAT. If anything it proves more data centres are needed. That's literally the only reasonable conclusion from this news.
- lifeisstillgood 2mo agoSorry I thought that a bubble was widely accepted. Are you arguing there is not an AI bubble, and that all the DC buildout is fine, going to be profitable etc? I am not looking for a online slanging match - just looking for a different point of view
- k2xl 2mo ago[dead]
- deyiao 2mo ago[dead]
- readthenotes1 2mo agoI wonder if Erdos would be saying " It's fine that y'all are answering my questions, but who is asking better questions??"
- robinhouston 2mo agoIn a way the most remarkable thing about this is that it isn't even at the top of the HN homepage. Even if this is a step up from what we've seen before, we're no longer astonished by the idea that AI can make significant advances in mathematics and computer science.
- schleck8 2mo agoThis is one of the most impactful mathematical publications in history by all accounts I think we've now hit a point where 99.9% of the population gloss over these types of AI advancements because of human competence being insufficient No human could have published this because it requires paradigm shifts (e. g. Section 5) in multiple mathematical domains. Mastering one of them to this degree is rare, mastering 3+ pretty much non existent for humans.
- antirez 2mo agoThis is not at the top as it is actively flagged by people that can't psychologically cope with the advances of AI. Hacker News is no longer a web site of an elite.
- deleted 2mo ago[deleted]
- pistoriusp 2mo agoInteresting. I had no idea that a person could see what is flagged?
- defrost 2mo agoIf you page through the /newest listings you can see [flagged] and [flagged][dead] submissions. eg. this: [flagged] A migrant surge tests Spain's open policies (economist.com) - https://news.ycombinator.com/item?id=49131860 https://news.ycombinator.com/item?id=49131860 is clearly marked as flagged. Unlike the current submission: Ten advances in mathematics and theoretical computer science (openai.com) which isn't [flagged]. * https://news.ycombinator.com/newest https://news.ycombinator.com/newest
- melagonster 2mo agoWow, so this is the end of science :(
- xyzsparetimexyz 2mo agoIt's just another tool that can help solve problems. It doesn't know _what_ problems to solve. It turns out that a lot of old problems are now low hanging fruit for these new models. In terms of 'expanding the frontier', we've just discovered dynamite and can now blast our way through mountains. The bottom of the ocean or space are still as hard to reach as ever.
- silver_sun 2mo agoIt's not even predictable like dynamite. Sometimes it can blast through a mountain, impressively, the problem is you can't predict which mountain it works on. And other times it can't even make a dent in a molehill, which is perplexing given what it was capable of earlier. Can we even call it dynamite?
- woeirua 2mo agoNo bud, it’s just the beginning!
- AngryData 2mo agoIt solved a handful of novel esoteric problems out of hundreds fed to it. Far from the end of science.
- DrBazza 2mo agoReplace philosophers for mathematicians and Douglas Adams was spot on again. Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this. -- "Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!" "What's the problem?" said Lunkwill. "I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!" "We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!" "You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"
- pama 2mo ago> Whilst current models can't 'intuit' and come up with conjectures I disagree. I routinely let LLMs speculate or generate hypotheses along the way of helping with technical research. Sometimes they can prove the correctness of a concrete math idea but other times even an unproven conjecture helps with the numerical algorithm implementation and the result is then simply supported by additional data. I guess that any autoresearch-adjacent application has LLMs intuiting and coming up with hypotheses/conjectures—as do the steps/lemmas along a complex proof. In my opinion the modern LLMs are powerful intuitive thinkers that generate lots of conjectures of varying quality or importance.
- evenhash 2mo ago> Whilst current models can't 'intuit' and come up with conjectures People keep saying this. Why? Surely the AI can complete the prompt “Generate new research questions based on these observations”? When I read the reasoning traces of coding models they are constantly asking themselves questions and attempting to answer them.
- kingstnap 2mo agoIt's remarkable how you can manage to get these models to produce remarkable breakthroughs like an explicit construction of a non-sofic group. And yet this is the exact same company that has screwed up their android app so bad that the latex N^3 rendering problem makes it so having it explain it to me crashes the app. Truly jagged beyond belief.
- xyzsparetimexyz 2mo agoAny implication of any of these findings? They seem like unimportant nerd snipes to me. If you want to do something actually relevant, get chatgpt to write a simulation of graphene nanotube construction and figure out how to do it at scale.
- utopiah 2mo agoVery marketable nerd snipes indeed.
- foobar10000 2mo agoOne - and I do not mean to be snarky - you can literally ask Gpt 5.6 Sol this - and if you want to see cool stuff - Fable running in their app (not website) has a view thinking button that is actually a good way to explore the adjacent fields, etc. The non-sofic group one is definitely a big deal - would have been a Fields medal if discovered by a human.
- Chance-Device 2mo agoPretty cool. The impact of AI is getting undeniable, there aren’t many positions left to move the goalposts to at this stage, next they’ll have to be outside the stadium entirely. The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
- datakan 2mo ago[flagged]
- Chance-Device 2mo agoI can deal with apathy, that’s the norm. What bothers me are all the people who think they can suppress AI by talking it down. That’s what’s counterproductive, just pretend the problem doesn’t exist. Tell other people it doesn’t exist either. I get it, it’s threatening socially, economically, maybe existentially. It’s also not going away.
- ryan_n 2mo agoSo you think it’s a potentially existential threat but are bothered by people who maybe want to suppress it… Hopefully you acknowledge there is a bit of lack of self awareness here eh?
- Chance-Device 2mo agoYou seem to have misunderstood my point.
- FranzFerdiNaN 2mo agoIt’s not apathy. It’s the fact that almost nobody can really understand what these results mean. I’m not a mathematician so I have zero clue what “ New upper bounds on sphere-packing density down to the Cohn–Elkies thresholds” means.
- danparsonson 2mo ago
- deleted 2mo ago[deleted]
- artninja1988 2mo agoNow that we've seen AI produce a fair number of proofs (and disproofs), I'm curious when we'll start seeing it build genuinely novel theory. Does anyone have predictions on when and how we'll get there and will it take new architectures/ training paradigms, or is the current approach enough?
- Davidzheng 2mo agoThere's no clean line between a collection of theorems and a theory.
- artninja1988 2mo agoI mean doing something like Grothendieck when he redeemed algebraic geometry or Galois when he invented group theory. We haven't seen that at all from LLMs.
- laichzeit0 2mo agoI’m personally hoping for the next big AI gangbanger to be theoretical physics. Boy does that field need a good reshuffle. I think when any novel mathematical theory can be done by AI you’ll see simultaneously theoretical physics getting wrecked as hard as pure math is. At that point we might see new physics or paradigm shifting technology emerging.
- slashdave 2mo agoIt will not happen with existing LLM techniques.
- bifftastic 2mo agoAny advances in theoretical physics yet? Are there any fundamental obstacles? I would have thought not, but I haven't seen anything reported.
- QuesnayJr 2mo agoThe Maxwell conjecture was a conjecture in theoretical physics (though not a particularly important one)
- ls612 2mo agoThe fundamental obstacle is that we have no conceivable way to produce the energy levels to test the predictions that new theoretical physics would produce. We are like over a dozen orders of magnitude off.
- tim333 2mo agoThere's a lot of everyday stuff in physics which is unexplained like the particle masses we have.
- Windchaser 2mo agoAnd a lot of condensed matter physics. Type II superconductivity is a well-known one, but there are a lot of more less well-known ones
- ultimatefan1 2mo agoone of the early premises of how ai takeoff would go was that a system that could solve open problems in advanced mathematics would also discover novel advances in math and computer science that directly unlock drastically better software performance. we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B). we are also seeing incredible advances in software performance. open ai announced like 15% improvement by fixing gpu kernel issues. these are clearly linked in the sense of scaling laws and generalization of intelligence: a huge model gets capabilities in both math and software engineering that isn't possible at smaller scales. but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence. to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math (i also posted this on twitter @mlipman13)
- woeirua 2mo agoThis makes no sense. To believe this you have to think that the models are somehow being overfit explicitly on academic mathematics and it doesn’t carry over at all to more practical software engineering. I wouldn’t make that bet.
- threatofrain 2mo agoThis also makes the assumption that frontier math has all the long hanging fruits already taken... also very dubious.
- Ar-Curunir 2mo agoSome of the problems solved here, at least in CS, have been open for decades, and have been worked on by very smart leading researchers in the field, including Turing Award winners. Like, these would be best-paper awards at many top CS conferences.
- deleted 2mo ago[deleted]
- christofosho 2mo agoI would love more time and money put into real-world problems by these companies. Climate, food insecurity, pollution, technology for convenience and/or to help people have a higher quality of life. I'm sure they must do some of this type of work, right?
- braneloop 2mo agoYes, but all of those are orders of magnitude harder than math.
- amazingamazing 2mo agoThey are political problems, a computer could never solve them.
- adroitboss 2mo agoTell that to game theory.
- throwaway198846 2mo agoA computer could solve them by creating the right technological ,social, rhetorical and economical solutions but that would lots of money anyway
- amazingamazing 2mo agoWe already know the solutions
- christofosho 2mo agoI suppose it depends on which types of problems you're targeting. There is a lot of physical science, and theoretical science, that has gaps because there aren't enough people working on tooling to assist in things like calculation, generation, simulation, etc. I agree, some of the problems are more difficult. I don't think that's the case for all of them. And, besides, these companies could be demonstrating how to approach problems and where their users could spend tokens to help with these problems. Should not these companies try to work on these problems _because_ they are difficult?
- amazingamazing 2mo agoCan’t wait for this stuff to have quality of life increases for the average person. So far all I see is that AI has made owning a computer more expensive, made some jobs redundant, increased spam and distrust with questionable authenticity of content and of course made some Americans very rich.
- jetsetk 2mo agoDownvoters mind to explain?
- evenhash 2mo agoNot everyone works for Evil Corp. I work in the public sector and my work supports public health and safety initiatives. AI has allowed my team to get much more done than we would have otherwise which improves the quality of life of the people in my community. So I would like to counter your cynicism with a “YMMV” depending on who you work for.
- galleywest200 2mo agoExamples?
- BoggleOhYeah 2mo agoHN doesn’t like when you ask for those. Remember, this is the propaganda arm of a VC. Proof is a hindrance to the hype that drives value.
- frenzyguy 2mo agoThis is both awesome and terrifying for mathematicians, however some ideas can be generated and the field as whole expanded with the attention! However, I was looking at the proofs and reason explanation and openAI should be more explicit in how the work has flown. I find the models have jumped hoops in some places of the proofs, that can be hard to track. In fact, when a paper is published you usually get a review and if no reviewer understands they ask you to further explain the thought process. It will be fun to see if this happens here.
- kcexn 2mo agoNot being an expert in any of the fields OpenAI has "advanced" I don't want to prematurely downplay the significance of this contribution. However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing. It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus, or are they simply an effective method to exhaustively search the literature for the right combination of existing tools to apply to the problem? Essentially, did these problems seem like they had an intuitive answer and were feasible to prove before, just not high enough value targets for an expert to invest time into? Or were they fundamentally difficult prior to this point and it appears that AI has done something more than just throw the problem into a big solver.
- simianwords 2mo ago> However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing Your worry.... is because they used the word advanced? For marketing? The word is used very appropriately here. There were PhD's who spent a big part of their career tackling these problems.
- kcexn 2mo agoI have no idea how many PhD's have spent how much time of their careers tackling these very specific problems, and I doubt you do either. I'm trying to understand if these specific problems were the kinds of problems that would have justified an expert investing weeks or months to solve. Or if they were the kinds of problems that would normally have been given to students to investigate.
- hollowcelery 2mo agoThey are significant problems which experts have spent months or years studying. I heard a mathematician say that resolving non-sofic groups and Connes's rigidity would be career-defining for a mathematician.
- 2mo ago
- sashank_1509 2mo ago[flagged]
- unknownian 2mo agoYou shouldn't be getting downvoted for something that a majority of pure math and art enthusiasts believe to be true. The truth is many of these entrepreneurs and VCs are obsessed with AI not for money or human progress, but because it makes them feel closer to being a "god" rather than a mere mortal. Much of it (especially AI art) is out of spite for human creativity, which is done by mortals with limitations.
- eadwu 2mo agoTaking the stance of moral superiority is kind of funny. And pure math and art enthusiasts don't think they are closer to being a "god" from understanding/"discovering" math? Stop coping and deluding yourself mate. To begin with, whether AI is the one doing the discovering or not makes no difference. Any "pure math" person would aim to understand regardless - and would be quite glad that they have a longer paved path. Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism).
- unknownian 2mo agoLmao what a ridiculous response. Yes, some mathematicians and artists are in it to feel smart. But the vast majority also just enjoy the process. Having a computer do all the work for you and just typing prompts in ruins that completely. As Ronny Chieng said in his Harvard speech, the journey is the point. >Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism) Ignoring that I meant pure as in non applied math, let's just make it clear: you agree that mathematicians who are against capitalism encroaching on this process should be allowed to dislike it without criticism of being pretentious?
- AlexeyBelov 2mo ago> Stop coping Isn't coping a good and useful mechanism?
- macleginn 2mo agoI am duly impressed by the powerl of the nameless internal AI, but not a single human contributor's name listed anywhere? Did someone at least make this model a coffee?
- drdrey 2mo ago> The results were achieved by an internal version of Astra, our next major model.
- zogomoox 2mo agosurely some human regularly typed "think deeper, make no mistakes".
- zardo 2mo agoMy grandmother is very sick and the doctors need this proof to help her.
- nefarious_ends 2mo agolol I used a prompt like this one time when chatgpt was refusing to translate a snippet of japanese text. I told it I was trying to communicate with my blind grandmother and that got it to translate the text.
- maxutility 2mo agoNew advances in sphere packing? Let’s make sure AI doesn’t inadvertently engineer ice-9.
- Ey7NFZ3P0nzAe 2mo agohttps://en.wikipedia.org/wiki/Ice-nine https://en.wikipedia.org/wiki/Ice-nine
- scuppernong 2mo agothe people who crow in the comments of each of these posts about AI advances making human beings useless seem to bizarrely identify themselves with the AI, but none of them seem to have had any hand in building this technology. at best, they're power users. pure ressentiment.
- ltitu 2mo agoSo they are bribing 100,000 researchers with free accounts to work on their future unemployment.
- petilon 2mo agoAt what point can we say AGI has been achieved? What is the test? AI is solving mathematical problems that humans have not been able to solve for decades. Is that not enough? Sam Altman has said "If superintelligence can't discover novel physics, I don't think it's a superintelligence." Is that the test? How far away are we from AI discovering novel physics? It seems within reach.
- antonvs 2mo agoIt’s artificial, it’s general, and it’s intelligence. The people who believe “AGI” is an important and unattained goal need to start coining and defining their terms better.
- tim333 2mo agoIt depends on your definition. For me it would have to be able to do the stuff humans can do like make a cup of coffee (Wozniak test). Just maths isn't really general enough for the G in AGI.
- petilon 2mo agoIt can give you detailed instructions for making a coffee. Is that not enough? Actually making coffee requires more than intelligence, it requires eyes and limbs (i.e., robotics). Think about a human that is blind and does not have limbs. Does he not have natural general intelligence, even though he is not able to make a cup of coffee?
- amai 2mo agoHave blog posts replaced peer-reviewed academic papers when it comes to publishing advanced in science?
- heaney-555 2mo agoThese is mostly mathematics, not science, and they link the paper in the post: https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29ad73/unit-distance-proof.pdf https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29a... Mathematicians will tear it to pieces if any of it is fake!
- beering 2mo agoIf you were a mathematician and came up with any of these results, people would pay attention even if you published it on your blog. What’s the requirement for needing to publish it in a journal? OpenAI is not trying to achieve tenure.
- simonw 2mo agoThe GitHub repo with the Lean formalizations just came out a couple of hours ago: https://github.com/openai/ten-proofs https://github.com/openai/ten-proofs It also links to a paper written by an LLM where the model "reconstructs how the proof came together" based on the unpublished reasoning traces: https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf I wish they'd publish the prompts though!
- fooker 2mo agoExact prompts haven't mattered for about a year now.
- Alifatisk 2mo agoCare to elaborate? Curious about this. Is this because LLMs have been geared towards understanding user user intent behind a prompt rather than following the instructions exactly?
- fooker 2mo agoThere's a full fledged 'reasoning' step that basically expands your prompt. As long as you are not missing important information, how you word the prompt does not have any effect.
- Alifatisk 2mo agoOh yeah, I suspected it was something like this. Thanks!
- s4i 2mo agoIsn’t that a huge simplification? Of course the way you phrase the prompt can carry semantic meaning, maybe subtly, but still. And sometimes that matters a little and sometimes a lot. I’ve stopped numerous agent sessions over the last few weeks to reword my initial prompt to get the agent off an unintended track.
- casey2 2mo agoPeople weren't their strongest even when most did manual labor. Now that humans are free from mental labor we work on creating and optimizing the best exercises for each mind. Couple that with restructuring transport infrastructure and diets many people will be smarter and fitter than at any time in history. They won't be able to outrun an automobile or out think an autointelligence.
- cindyllm 2mo ago[dead]
- bwestergard 2mo ago"People weren't their strongest even when most did manual labor." Is there good historical data on some measure of strength across representative populations over time in the modern era? I'm doubtful. We do know that the introduction of agriculture diminished strength: "Bone mass was around 20% higher in the foragers - the equivalent to what an average person would lose after three months of weightlessness in space. After ruling out diet differences and changes in body size as possible causes, researchers have concluded that reductions in physical activity are the root cause of degradation in human bone strength across millennia." cam.ac.uk/research/news/hunter-gatherer-past-shows-our-fragile-bones-result-from-physical-inactivity-since-invention-of
- qnleigh 2mo agoCan anyone comment on the significance of any of these results for their respective fields? Or what impact they might have? Presumably none are quite at the level of the Jacobian conjecture, but some of the results on group theory and sphere packing sound pretty important at first glance.
- qnleigh 2mo agoFound some discussion here [1] from someone who actually worked on a few of these problems. [1] https://x.com/henryquantum/status/2083623695436623915?s=20 https://x.com/henryquantum/status/2083623695436623915?s=20
- Kotlopou 2mo agoCame here looking for this, sadly there's little engagement with the actual math in this thread. I hold out hope that Scott Aaronson might comment on the circuit bound eventually...
- MinimalAction 2mo agoI hate this timeline. I might be excited for the kind of answers this AI builds for unsolved problems, and also for learning new things by talking to it. But, I feel like I'm in the minority of people here who feel this could be a net negative endeavor with this having to kill a lot of educational institutions and their ability to fund themselves in the long run. It's not worth that.
- randomizedalgs 2mo agoAfter skimming some of the writeups, I'm surprised that the frontier internal model still writes just as poorly as Sol. Maybe good AI paper writing is further away than I thought...
- QwenGlazer9000 2mo agoYou mean we're still gonna be employed doing the boring part while AI gets to do the fun part? I'd honestly rather they just automate every job at that point.
- gpm 2mo agoHenry Yuen's (whose work problem 6 builds on) comments on this are worth reading IMO: https://bsky.app/profile/henryyuen.bsky.social/post/3ms2jpchfjc2t https://bsky.app/profile/henryyuen.bsky.social/post/3ms2jpch...
- an0malous 2mo agoIt sounds like he hasn't verified the results of a problem that he has personally worked on, so how many of these problems have actually been verified?
- gpm 2mo agoI mean, they're verified in the sense that the lean proof checks out... and presumably OpenAI read them.
- deleted 2mo ago[deleted]
- doctorwho42 2mo agoOr they made another LLM 'read' them? > You are an expert in the field of mathematics, with decades of experience. You are a reviewer of proofs, etc etc.etc.
- margorczynski 2mo agoFrom what I understand all of them have Lean proofs/certificates thus are basically 100% proven without a doubt.
- voxl 2mo agoIncorrect. The statement in Lean can itself be wrong. Moreover, they could be exploiting a kernel bug in Lean, of which we had one published literally a week ago.
- samrus 2mo agoWe recently saw that lean itself isnt proven correct. Its not likely but i wouldnt call it verified if its only verified in lean https://x.com/gro_tsen/status/2082483878480977959 https://x.com/gro_tsen/status/2082483878480977959
- sf12sd 2mo agoNot peer reviewed, Lean proofs are 100,000 lines long and Lean has bugs: https://cr.yp.to/proofs.html https://cr.yp.to/proofs.html Who is going to wade through this?
- kypro 2mo agoThey've been hiring mathematicians to verify this stuff themselves. They're obviously not just throwing it out there without any human review.
- 12asg 2mo agoAnd these mathematicians sink the comment to the bottom in 5 min? It is not peer review if it is all in one company that wants an IPO.
- drcongo 2mo agoThis thread has an absolutely wild points to comments ratio.
- big_toast 2mo agotomhow explains they gave the story another shot here: https://news.ycombinator.com/item?id=49158443 https://news.ycombinator.com/item?id=49158443
- Kelteseth 2mo agoWhat's up with the upvote/comments ratio 8 to 337 on this post? Are the comments already also ai advanced? (/s?)
- jsnell 2mo agoOriginal submission (460 votes) two days ago: https://news.ycombinator.com/item?id=49132058 https://news.ycombinator.com/item?id=49132058 For some reason comments got moved to this one.
- tomhow 2mo agoIts visibility was diminished due to the flamewar detector and most of its front page time being during overnight hours on Friday night/Saturday morning USA time. I've created a new copy to give it some primetime exposure, because it seems like an important enough announcement to warrant it.
- muchmirulys 2mo agoproblem number 1 and 9 are surprisingly very intuitive check here : 1. high dimensional sphere packing https://muchmirul.github.io/conjectures/sphere-packing/ https://muchmirul.github.io/conjectures/sphere-packing/ 2. multicolor ramsey number https://muchmirul.github.io/conjectures/multicolor-ramsey https://muchmirul.github.io/conjectures/multicolor-ramsey
- dash2 2mo agoThe first link is very sloppy and doesn't actually explain why the "certificate" proves anything about the sphere packing. Or if it did, I couldn't understand it.
- CGMthrowaway 2mo ago[dead]
- rothos 2mo agoAgreed
- kypro 2mo agoI want to iterate the most important thing about this is that it's yet more evidence of AI's accelerating competency in solving math and comp sci problems, and suggests we're now getting close to the point where you could throw AI at AI research challenges (which are largely just math and comp sci problems) and potentially find very real algorithm improvements. AI development is likely to be more compute bottlenecked than solving math problems since validation of any algorithmic improvement would likely require significant compute. But you could imagine that at this point it could be economical for a frontier lab to task 10,000 agents to work non-stop on finding novel algorithmic improvements then validating the top 50 out of 1,000 candidates on a GPT-2 sized network. I would suggest RSI is now very close. The singularity could be less than 6 months away. I'm not saying I'd put a high probability on that, but I'd give it at least 20%, and I'd double that if looking 12 months out. I know I'm just a crazy man shouting at the clouds, but please take to the consequences of this seriously. I understand that for whatever reason AI risk seems abstract and doesn't seem real, but this should terrify any person thinking logically about where this could all be heading. We haven't even solved the most basic AI safety problems yet. RSI right now would almost certainly result in an extremely bad outcome for humanity.
- xpct 2mo agoOkay, let's take it seriously. What do you propose? What can your average person do to prepare for RSI beyond bracing themselves mentally?
- reducesuffering 2mo agoYou can not prepare or brace yourself mentally any more than you can a terminal cancer diagnosis. An RSI foom right now means an unaligned superintelligence will disregard us in pursuit of its goals. We would be ants in the way of a data center being constructed. All people can do is collectively support the notion, like 1200+ frontier AI researchers and their CEOs, that we do not have control of where this is headed, we need to immediately slow down the race, in time for people to agree that we do not have the capability to align a superintelligence to humanity’s wishes
- 2mo ago
- deleted 2mo ago[deleted]
- maxprimes 2mo agoI'm sure OpenAI is just interested in the greater good of mankind!
- bryan0 2mo agoI think there's another interesting story here about how this was apparently moderately flagged and triggered the flame-war detector which kept the story off the front page of HN 2 days ago[0]. I think people are having a hard time processing this information rationally(?) What can we do to make conversations around these incredibly exciting and important topics more constructive? HN is where I expect to read expert comments on these topics, has this style of conversation moved elsewhere? [0]: https://news.ycombinator.com/item?id=49157930#49132926 https://news.ycombinator.com/item?id=49157930#49132926
- jrflo 2mo agoI think that HN is particularly negative towards AI because the vast majority of users here will have their prestigious CS careers disrupted by AI advances. So, there's an inherent negative bias towards this news. I for one am really fascinated by AI's advances in science and math and would like to talk about it somewhere without the constant flamewars...
- sothatsit 2mo agoExtreme claims on posts like these also, rightfully, trigger people’s skepticism. I don’t think it’s wrong to question claims that math is dead as a field. But then it leads people to miss the overall trendline. People argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps leading to crazier and crazier results. The much more interesting question to me is what will be consumed by the exponential like math seems to be, and what won’t. Writing has been much more stubborn, but I’ve noticed Fable to be quite a big step up there as well. How about politics? Will we develop new ways to let people express their own values in democracies, or will we get much better at manipulation? And then there’s questions like, even if AI can answer increasingly complicated math questions, will we still need mathematicians to translate results to the real world, verify them, or decide where to push the frontier?
- tomhow 2mo agoThe main issue was that it was submitted late Friday night SF time, meaning it was overnight or Saturday everywhere in the world when the post had its chance on the front page. The flagging was minimal relative to the vote count and had no effect, and the flamewar detector would have been turned off sooner if moderators saw it sooner (it wasn't really a flamewar, just a lot of comments). It still spent 10 hours on the front page. None of this is anything out of the ordinary; this kind of thing has always happened. The only real story here is that moderators sleep sometimes.
- overgard 2mo agoGary Marcus' has a good take on this: https://garymarcus.substack.com/p/openais-amazing-but-vastly-oversold?r=8tdk6&utm_campaign=post-expanded-share&utm_medium=web&triedRedirect=true https://garymarcus.substack.com/p/openais-amazing-but-vastly... https://garymarcus.substack.com/p/two-critical-updates-re-astra-and?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4376bf6-f1ba-4471-88b5-4307e2581a40_1213x1027.png&open=false https://garymarcus.substack.com/p/two-critical-updates-re-as... Not that there isn't something interesting in here, but lets be clear that we don't have enough information to evaluate this properly. And as always with these labs, BS takes a lot more energy to refute than it does to spread.
- HardCodedBias 2mo ago"Gary Marcus' has a good take " I think that is an oxymoron.
- neta1337 2mo agoHow so? His predictions were accurate so far
- energy123 2mo agoNo they were not. These were his 5 predictions in 2022: """ 1. By 2029, AI will still be unable to watch a movie and accurately explain the characters, events, conflicts, and motivations. 2. By 2029, AI will still be unable to read a novel and reliably answer questions about its plot, characters, conflicts, and motivations beyond what is stated literally. 3. By 2029, AI will still be unable to work as a competent cook in an unfamiliar kitchen. 4. By 2029, AI will still be unable to reliably create more than 10,000 lines of bug-free code from natural-language instructions or interaction with a nontechnical user, excluding simple assembly of existing libraries. 5. By 2029, AI will still be unable to convert arbitrary mathematical proofs written in natural language into symbolic form suitable for formal verification. """ There's still 3 years to go and he's already wrong on 4 out of 5.
- 2mo ago
- merelydev 2mo agoGreat stuff. Wonder how many of the ten problems where solved by independent mathematicians not linked to OpenAI
- p1esk 2mo agoZero. These were open problems.
- dipanshuhappy 2mo agoCrazy progress. I wonder how institutional academia would adjust with this. Now its more apparent than ever that the prestige and honour system in academia is having shaky foundations
- titanix88 2mo agoHow do we know that these solutions don't exist in the training data? It is open secret that they have used pirated materials for training. Perhaps it plagiarized solutions from works of some obscure Belgian mathematician from the sixties, who did not get mainstream acceptance. I wouldn't be surprised if they also got access to mathematics done in the "defense contractor" setting from various three letter agencies. Without a searchable index of training data, it is hard to put faith into these claims.
- ken47 2mo agoThis wouldn't be a problem so long as they properly attribute.
- jgord 2mo agoupvoted for fair point .. its possible an LLM AI could be put to work to search widely for attribution / similar results. eg. "we spent another 2k on searching for pre-existing proof but found only the weaker result xyz by abc in 1972" would be in the spirit of academics quoting prior work.
- cwiz 2mo agoI feel increasingly anxious reading this. Machine research shouldn’t be merged into mainline of human knowledge.
- MattGaiser 2mo agoKnowledge is knowledge, as long as it can be proven true.
- raver1975 2mo agoproving false is also useful
- vessenes 2mo agoWhen you can formalize it in Lean or some such, why would this be? I can understand the desire to separate out other forms of research from the human corpus. But theoretical math that is decidable/provable, I’m not sure I see the risks.
- rencrisa 2mo agoI just want to state that having "lean proofs" that build does not mean the actual real theorems we care about hold. Ultimately a human has to verify the lean encoded theorem statements that the lean proofs are checked against. For non-trivial theorems such as these, this is an arduous and tricky task where even a little mistake could be fatal.
- vessenes 2mo agoAgreed that confirming “This proposition has been faithfully translated to Lean” matters. As a side note, though, this is an area where LLMs used interactively can be helpful. Also in the case of these recent counterexamples turning up, we have good qualitative evidence that this pathway is effective.
- jstummbillig 2mo agoWhy?
- bhawika_kaushik 2mo ago[dead]
- areoform 2mo agoLooking at this thread, I can see that a lot of technical people have ambivalent to negative feelings towards AI, but with each new generation, I become more and more convinced that they're missing out on something interesting. It is indeed true that all models are, at their core, predictors of what occurs next in a sequence. But I think it's worth exploring the implication of what that means. Because when fed tiny pieces of information for a few tasks at a small scale, this results in something that sorta, kinda works. Or, works surprisingly well. But when scaled... When the amount of information starts approaching the sum of all human knowledge, the tasks start approaching all useful applications of that human knowledge, and the fidelity of the predictor approaches incomprehensible sizes, the starts encodes / becomes (I'd argue it becomes) something that can model all human knowledge. It feels wrong to say that, but let me explain, what is the best way to predict the behavior of a ball constrained in two directions that bounces with initial vertical velocity v(y) (y is up / down axis) and horizontal velocity v(x) (x is side by side in 1d) ? If we purely look at it via a graph, it's by modelling the function of acceleration under earth's gravity. If only a few points are given to you for this and you can't make something really sophisticated, then you'll make something that's rough that kinda sorta works and then call it a day. But... if the number of points keeps increasing in number, precision and accuracy as well as the number of examples (assumed that data about air pressure, velocity and all other factors is included alongside these points), the fidelity with which you can replay / tweak the function keeps improving, and the number of times you can iterate keeps increasing, you'll eventually create a function that models that process so well that it intrinsically contains a good enough model of the deformation of the ball (provided the dataset contains information about elasticity of the ball's material, its dimensions and mass etc..), the nearly negligible (under normal conditions) effects of the ambient environment (provided there's diversity in the number of environments supplied), the oblateness of the Earth and minute changes in the gravitational field (the length of a seconds pendulum varies depending on where the experiment happens. It's presumed that all of the prior set of experiments were repeated across the Earth and the subtle, but real deviations were faithfully recorded)... and so much more. A machine trained on the above with a large number of parameters, measures to prevent "laziness" and enough reps for high fidelity across a large enough dataset would start to approach a simulation of the ball falling. Because to predict what happens next in the sequence, you must model what's occurring in the sequence. Now imagine doing that for other tangible and intangible things in this world. For all of human knowledge across all fields of endeavor. All experiences. No matter how noble, ignoble, notable or ignorable. But putting all of it into the soup that's this machine. Then at larger and larger scales, you eventually start encountering "good enough" models (in modelling the falling ball sense) for even the most hard to quantify / qualify things like grief and joy. At some point, by simply trying to predict what it has been taught ought to be the next part of the sequence in say... human interaction, it starts to make a model of something that hews ever closer to a full fidelity theory of mind. Is there evidence for this? Kind of, yes. There are early indications that as machines are trained for an ever larger number of tasks at larger and larger scales, their internal representations converge. It's called the Platonic Representation Hypothesis. Overview and paper here, https://phillipi.github.io/prh/ https://phillipi.github.io/prh/ It is my opinion that these machines are displaying a new form of intelligence that human beings haven't quite encountered before. They are the sum of all human knowledge made manifest and given voice by processes that nudge (bit-by-bit) what kind of step it ought to predict for the next part of whatever sequence it displays. In my mind this means that, of course, these models can create new knowledge. This strains the analogy, but with the sum of all human mathematics within them, they can "reason" via the act of predicting what ought to come next. Of course, these machines are "surprisingly" good at a lot of things the larger they get, because what the labs have created here is a rough version of humanity's collective knowledge given form and the ability to say hello. I suspect that the current generation isn't close to the "true frontier" of what these machines could be. They are nowhere close to the sum of all human knowledge and endeavor. They are quite a way there, but they haven't yet achieved true completeness for domains where the data isn't so public. I think it's the most exciting scientific and technological breakthrough of my lifetime. And I can't wait for us to get close to the true frontier of all domains.
- sothatsit 2mo agoPeople argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps delivering more and more impressive results. The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn, but I’ve noticed Fable to be quite a big step up there. How about politics? Will we develop new ways to let people express their own values in democracies, or will we just get much better at manipulation? How about experiment driven domains like biology?
- jcims 2mo ago>Will we develop new ways to let people express their own values in democracies, or will we get much better at manipulation? Yes.
- tyre 2mo agoWe will get much better at manipulation and better at people “writing” things to justify their own feelings. What’s new about LLMs is that you can scalably manipulate people individually. It used to be that you could either have scale (speeches, tweets, interviews, website, etc.) or individual engagement (replying to mail/tweets/town hall questions.) Now you can pull the history and preferences of an individual, then shape a message—in real time—to them, specifically. You can have conversations on social media with a single person and shape your message specifically to them. Part of this can be good (you talk about what they care about, where 90% of broadcast messaging might not apply) and part of it can be bad (manipulation.) My guess is that, in the US, the right will cynically adopt manipulation to great effect and the left will take a moral stand against shady practices and lose elections.
- lettergram 2mo ago> My guess is that, in the US, the right will cynically adopt manipulation to great effect and the left will take a moral stand against shady practices and lose elections. I think that statement may itself highlight how prevalent manipulation is. I fully anticipate all groups to continue maximal manipulation they can. One thing with LLMs is that it'll be a far less unified view, so a "divide and conquer" strategy is what I anticipate.
- raver1975 2mo agoI wish I could qualify for some free AI as a mathematics researcher. I guess I'm just an amateur. https://alethean.org https://alethean.org
- pbkompasz 2mo agoWow, here are solutions to 10 problems that we spent millions of dollars on out of the 100s/1000s of other problems that we tried and failed to solve.
- deleted 2mo ago[deleted]
- namr2000 2mo agoI understand the frustration with the constant PR-hype these AI labs keep spewing out, but on other hand I just can't understand this sentiment at all. These are real problems mathematicians and computer scientists have been working on and were unable to make progress on. Now they have been given a new tool and using that tool have solved those problems. And its not just one or two problems, its many very difficult problems. The mathematicians I know are saying that the latest crop of models is changing the way people do research math, I think that's a pretty big deal.
- class3shock 2mo ago"I understand the frustration with the constant PR-hype these AI labs keep spewing out" Apparently you don't. "These are real problems mathematicians and computer scientists have been working on and were unable to make progress on." Who says no one was making progress? Who says openai has made progress? How would anyone not working on these specific problems, witho the time to dig into openai's claims, be able to tell? Why should this not be lumped in with all the other ai hype being pushed? "The mathematicians I know are saying that the latest crop of models is changing the way people do research math, I think that's a pretty big deal." Who? And doing what? We have been hearing the "this generation of models is the one" type talk for years and the only concrete "big deals" are what? A tool for college students to write papers? A replacement for, now enshitified, google search? The fact that now you can fake tons of stuff to support a position or claim tons of stuff that goes against your position is fake?
- namr2000 2mo agoI want to preface my response by saying that I don't buy most of what the AI labs say. I don't think that LLMs will replace most white collar labor for example. I also find many of the practices of these labs to be abhorrent. However, all of these opinions are orthogonal to the fact that LLMs have gotten extremely good at mathematics. > Who says no one was making progress? Let's look at the Jacobian conjecture, since that was the open math problem I was most familiar with prior to its solution. Yitang Zhang, one of the worlds most renown mathematicians (famous for his lower bound on the twin prime conjecture) spent 8 years working on this problem with his advisor (who himself is a renown mathematician) and turned up completely empty handed. His advisor described it as a "waste [of] 7 years of his own life and my time" [1]. Of course, these two were not the only ones working on this problem for the almost 100 years its been open, but they should have sufficient credentials to show that they were not fools or amateurs. And in a single afternoon an LLM disproved the conjecture. How is that not an extraordinary feat of technology? > Who? And doing what? A close friend is studying differential geometry in a PhD program. Sadly I doubt anything I say on his work will convince you, so I will instead offer two anecdotes: Terrence Tao (widely considered the worlds greatest living mathematician) has said AI is precipitating "a crisis in the foundations of mathematical values and practices" [2]. Timothy Growers (fields medalist & one of the leading researchers in combinatorics) has said that the latest models are now at the point where they are "producing a piece of PhD-level research in an hour or so, with no serious mathematical input from me" [3]. You can find many more fields medalists and mathematics researchers with the same impression. If you look in this thread you can see bluesky/twitter threads from those who were actively researching some of these problems who are in shock at the solutions. [1] https://www.math.purdue.edu/~ttm/ZhangYt.pdf https://www.math.purdue.edu/~ttm/ZhangYt.pdf [2] https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.p... [3] https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/ https://gowers.wordpress.com/2026/05/08/a-recent-experience-...
- 10dpd 2mo agoWhile these advances are genuinely impressive, I'm curious when we will see practical implications for this work. For example, will we see advances in material science, medical cures, etc? Would love to read about some examples of practical impact.
- oblio 2mo agoThe big example predates LLMs as a unified tech and it's protein folding, from Google DeepMind. OpenAI and Anthropic are too greedy for cash to do anything of the sort. I don't expect this current economic cycle to bring anything else that will directly greatly improve the life of the average person on the planet, more than it hurts it.
- QwenGlazer9000 2mo agoAnd that's the crux. LLMs are the nuclear bomb, while AI is the fission. People refuse to stop building them even if they make things worse.
- plaidfuji 2mo agoAny computable problem will eventually fall to computers. LLMs have made math proofs more computable, in the sense that a computer can both generate potential solutions and check the validity of its solutions on its own, with a reasonable chance of converging on something correct. I assume this was already doable to some extent, but it seems like it’s now exponentially easier. That still doesn’t mean that all math is automatically solved. This is somewhat similar to things like molecular dynamics or protein folding or finite element simulations, etc. Some problems that were previously intractable via computation became tractable. Others - the vast majority of other problems - remain unsolvable by these computational techniques, because the scale of compute required is beyond imagination. These are simple things like simulating the dynamics of a cubic millimeter of water molecules for 1 second. Unfathomably beyond current capabilities (and LLMs aren’t going to change that). I think LLMs are great, I use them every day and I think they have a ton of value. But if these things were as revolutionary as people promote/fear them to be, you should immediately point them at the highest value math problems and see progress. Like the Millenium Prize problems. Haven’t seen a solution to those. So there are limits - but we’re about to learn a lot about the new normal of what constitutes a layup math proof vs the truly difficult.
- VladVladikoff 2mo agoWould be great to see them solve Yang-Mills and Mass Gap.
- gerdesj 2mo ago"Any computable problem will eventually fall to computers." By definition. Its those pesky NP jobbies that get in the way.
- gerdesj 2mo agoJust to re-iterate the point: Whenever the handwaving starts around a discussion relating to a NP hard problem, I find it useful to imagine a Canadian bloke (MHRIP) in a red top, with a ... Scottish accent ... saying: "Ye cannae break the laws o' physics, Jim". (maffs not fisics, obvs!) If that is a bit tiresome for the gung-ho AI evangelist, there is also the rather knotty snag that that blasted Austrian geezer Gödel fiddled up: incompleteness. Its almost as though these bloody clever scientific and that types keep on putting artificial blocks in the way of LLMs laying golden eggs! I'm quite happy with the "marginal gains" I get with a DGX Spark. It will pay for itself within three months doing stuff on prem and us not sending data to someone else. It will scale.
- catching_crumbs 2mo agoOnly left is to catch crumbs from the table (2000): https://gwern.net/doc/fiction/science-fiction/ https://gwern.net/doc/fiction/science-fiction/
- tanh 2mo agoI feel nothing but grief.
- kart23 2mo agoAI can do this shit but can't do the dishes
- deleted 2mo ago[deleted]
- nnm 2mo agoThis is impressive. The real game-changer will be when AI creates an entirely new, significant branch of mathematics.
- andai 2mo agoWhich only it understands?
- energy123 2mo agoI would argue we're crossed that point recently. This is a comment from twitter that I appreciated: > "Already, there are very few mathematicians qualified to verify OpenAI’s new results. As progress continues, that number will approach zero." Not that I'm good enough at math to have any uniquely formed opinion, but after reading commentary from people who are, my impression is that these new results are bamboozling the humans due to using tools from so many disparate areas. At least, we can say that there isn't a single human who is smart enough to understand all ten proofs, even if there is a collective sense in which all proofs are understood.
- written-beyond 2mo agoInteresting how HN promoted this post to the front page again with a fake submission time. There are comments two days old, seems weird why they'd want this post specifically to get more traffic.
- john_strinlai 2mo agonot very interesting, its the second chance pool. its not a conspiracy of ycombinator vs. openai. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=second%20chance%20pool&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- oersted 2mo agoOf course that’s the mechanism and this kind of repost is not unique, but it is still relevant to note that the small HN committee made an exception for this one.
- mrloopex 2mo agoAs someone who’s spent the last year deploying PQC cryptosystems, it sure would be a kick in the ass if they find a faster solution to the nearest vector problem. I mean that’s why we’re going hybrid, but still. (Well to be more precise we’re going hybrid because of unknown SCAs.)
- c0rruptbytes 2mo agohttps://vibemathed.com/ https://vibemathed.com/
- navaed01 2mo agoIs there anyone here who can comment if these advances are meaningful, novel and how extraordinary these advances are? (E.g most PHDs, top 10% of professors etc. )
- throwaw12 2mo agoThis proves solving open problems in Mathematics were a search problem. But it might be not good for human brains, because we trained our brains with these problems and our brain optimized search space in some ways, and yes, we also couldn't solve some these problems. Now imagine someone gets stuck with a problem which could become its own theory, but they will solve it with LLMs and move on to solve their primary problem, because they don't realize how other problem was a big deal. If theory is not formalized, then it won't contribute to the search space for other person, solving different problem. All in all: * people's brain will be shaped differently * we will lose search space optimizations in our brains * we lose new theory contributions, which increases the search space to help solve other problems
- haowen8098 2mo ago[flagged]
- MichaelMoser123 2mo agoIt didn't come up with a counterexample for the P versus NP problem, I wonder if they just didn't ask about it...
- samuelknight 2mo agoYou can't come up with a counterexample for P != NP because there isn't a formula to disprove. For P = NP you would propose a general algorithm to convert all NP problems into P in P time, and an AI could then find a counterexample which would disprove that particular method. To demonstrate P != NP you need to prove that no possible algorithm can convert any NP into P which is much harder than providing a counterexample. AI has just gotten to the intelligence that it can make clever counterexamples to mathematical conjectures, but the frontier isn't quite smart enough that it can make novel contributions to mathematics. We are really close though. Only a matter of months away.
- isaacfrond 2mo agoThe paper does not resolve P versus NP, but it does make an important advance in a closely related area. To prove that P ≠ NP, it would be enough to show that every algorithm for an NP-complete problem requires superpolynomial time. We cannot prove anything remotely that strong. For explicit NP-complete problems in unrestricted computational models, we cannot even prove superlinear lower bounds. There is therefore an enormous gap between the lower bounds we can prove and the superpolynomial bounds we would need. VP and VNP are closely related algebraic analogues of P and NP. Here the paper proves new lower bounds for computing the permanent, a VNP-complete polynomial, in particular models of arithmetic computation: roughly (n^2\log\log n) arithmetic gates for unrestricted division-free circuits, and (n^4/\log n) size for the more restrictive formula model. These are still polynomial bounds, so they do not separate VP from VNP. But lower bounds on the resources needed to compute explicit functions are exactly what would ultimately be required for such a separation, and meaningful lower bounds of this kind are exceptionally rare.
- deviation 2mo agoGenuine question, from someone with a non-math background - at what point does the constraint become the number of unsolved conjectures remaining, instead of the ability for LLMs to actually solve for one?
- datsci_est_2015 2mo agoNot entirely following the question, but there are an infinite number of conjectures, the blinding majority of which serve nearly no purpose to humanity. Consider that for every executable program one could create a conjecture, and therefore a mapping exists from executable programs to conjectures. Now, consider the infinite possibility space of executable programs… Anyway, unless you mean conjectures that humans have already posited, or ones that are particularly famous, that list is much shorter, but also contains conjectures that I’m not convinced can be solved before the heat death of the universe using all available compute power. P != NP is a conjecture, for example. Also a lot of prime number conjectures that are extremely computationally expensive.
- Kotlopou 2mo agoThe oldest unanswered math problem known, based on (https://mathoverflow.net/questions/27075/what-is-the-oldest-open-problem-in-mathematics https://mathoverflow.net/questions/27075/what-is-the-oldest-...), is whether there are any odd perfect numbers (= numbers that are equal to the sum of their divisors). It's been open for 1900 years. Math is very hit-or-miss; the complexity of a question does not give much of an indication about how complex the answer will be. Look up the formula for solving a degree-4 polynomial equation to get a purely visual idea how this can look (and then degree-5 suddenly forces you to use complicated new functions). And there are problems (like the Collatz conjecture or P vs. NP) that there doesn't seem to be any promising angle of attack for over at least decades. I would wager that this is a fundamental part of the structure of math that has been a constant from ancient Greece till LLMs. There are even some formal results, similar to Gödel's theorems, that say that the maximum necessary length of a proof grows arbitrarily fast (e.g. more than exponentially, double-exponentially, or any function with a formula) with the length of the statement being proven. Point is, math will most likely never suffer from this particular problem.
- jimmyswisher 2mo agoI recently read Fermat’s enigma and I really loved it and am amazed at how far humans can go. If AI does everything going forward it will definitely be bitter sweet taking away from the human possibility of it imo