12 ms·
AlphaProof's Greatest Hits
- deleted 2y ago[deleted]
- wslh 2y agoIf you were to bet on solving problems like "P versus NP" using these technologies combined with human augmentation (or vice versa), what would be the provable time horizon for achieving such a solution? I think we should assume that the solution is also expressible in the current language of math/logic.
- hiddencost 2y agoNo one is focused on those. They're much more focused on more rote problems. You might find them used to accelerate research math by helping them with lemmas and checking for errors, and formalizing proofs. That seems realistic in the next couple of years.
- nybsjytm 2y agoThere are some AI guys like Christian Szegedy who predict that AI will be a "superhuman mathematician," solving problems like the Riemann hypothesis, by the end of 2026. I don't take it very seriously, but that kind of prognostication is definitely out there.
- titanomachy 2y agoI’m sure AI can “solve” the Riemann hypothesis already, since a human proved it and the proof is probably in its training data.
- nybsjytm 2y agoNo, nobody has proved it. Side point, there is no existing AI which can prove - for example - the Poincaré conjecture, even though that has already been proved. The details of the proof are far too dense for any present chatbot like ChatGPT to handle, and nothing like AlphaProof is able either since the scope of the proof is well out of the reach of Lean or any other formal theorem proving environment.
- Davidzheng 2y agowhat does this even mean? Surely an existing AI could reguritate all of Perelman's arxiv papers if we trained them to do that. Are you trying to make a case that the AI doesn't understand the proof it's giving? Because then I think there's no clear goal-line.
- nybsjytm 2y agoYou don't even need AI to regurgitate Perelman's papers, you can do that in three lines of python. What I meant is that there's no AI you can ask to explain the details of Perelman's proof. For example, if there's a lemma or a delicate point in a proof that you don't understand, you can't ask an AI to clarify it.
- Davidzheng 2y agolink to this prediction? The famous old prediction of Szegedy was IMO gold by 2026 and that one is basically confirmed right? I think 2027/2028 personally is a breakeven bet for superhuman mathematician.
- nybsjytm 2y agoI consider it unconfirmed until it happens! No idea where I saw it but it was probably on twitter.
- zarzavat 2y agoProbably a bad example, P vs NP is the most likely of the millennium problems to be unsolvable, so the answer may be "never". I'll bet the most technical open problems will be the ones to fall first. What AIs lack in creativity they make up for in ability to absorb a large quantity of technical concepts.
- wslh 2y agoThank you for the response. I have a follow-up question: Could these AIs contribute to advancements in resolving the P vs NP problem? I recall that the solution to Fermat’s Last Theorem relied on significant progress in elliptic curves. Could we now say that these AI systems might play a similar role in advancing our understanding of P vs NP?
- ccppurcell 2y agoJust my guess as a mathematician. But if LLMs are good for anything it will be for finding surprising connections and applying our existing tools in ways beyond human search. There's a huge space of tools and problems, and human intuition and brute force searching can only go so far. I can imagine that LLMs might start to find combinatorial proofs of topological theorems, maybe even novel theorems. Or vice versa. But I find it difficult to imagine them inventing new tools and objects that are really useful.
- nemonemo 2y agoDo you have a past example where this already-proven theorem/new tools/objects would have been only possible by human but not AI? Any such example would make your arguments much more approachable by non-mathematicians.
- meindnoch 2y agoOk, then the AI should formally prove that it's "unsolvable" (however you meant it).
- 2y ago
- uptownfunk 2y agoThe hard part is in the creation of new math to solve these problems not in the use of existing mathematics. So new objects (groups rings fields) etc have to be theorized, their properties understood, and then that new machinery used to crack the existing problems. I think we will get to a place (around 5 years) where AI will be able to solve these problems and create these new objects. I don’t think it’s one of technology I think it’s more financial. Meaning, there isn’t much money to be made doing this (try and justify it for yourself) and so the lack of focus here. I think this is a red herring and there is a gold mine in there some where but it will likely take someone with a lot of cash to fund it out of passion (Vlad Tenev / Harmonic, or Zuck and Meta AI, or the Google / AlphaProof guys) but in the big tech world, they are just a minnow project in a sea of competing initiatives. And so that leaves us at the mercy of open research, which if it is a compute bound problem, is one that may take 10-20 years to crack. I hope I see a solution to RH in my lifetime (and in language that I can understand)
- wslh 2y agoI understand that a group of motivated individuals, even without significant financial resources, could attempt to tackle these challenges, much like the way free and open-source software (FOSS) is developed. The key ingredients would be motivation and intelligence, as well as a shared passion for advancing mathematics and solving foundational problems.
- uptownfunk 2y agoOk but how do you get around needing a 10k or 100k h100 cluster
- wslh 2y agoIt is well known that cloud services like Google Cloud subsidizes some projects and we don't even know if in a few years improvements will arise.
- uptownfunk 2y ago
- sbierwagen 2y agoMore information about the language used in the proofs: https://en.wikipedia.org/wiki/Lean_(proof_assistant) https://en.wikipedia.org/wiki/Lean_(proof_assistant)
- sincerely 2y agoin the first question, why do they even specify ⌊n⌋ (and ⌊2n⌋ and so on) when n is an integer?
- rishicomplex 2y agoAlpha need not be an integer, we have to prove that it is
- sincerely 2y agoShould have read more carefully, thank you!
- Robotenomics 2y ago“Only 5/509 participants solved P6”
- nybsjytm 2y agoThis has to come with an asterisk, which is that participants had approximately 90 minutes to work on each problem while AlphaProof computed for three days for each of the ones it solved. Looking at this problem specifically, I think that many participants could have solved P6 without the time limit. (I think you should be very skeptical of anyone who hypes AlphaProof without mentioning this - which is not to suggest that there's nothing there to hype)
- auggierose 2y agoCertainly an interesting information that AlphaProof needed three days. But does it matter for evaluating the importance of this result? No.
- nybsjytm 2y agoI agree that the result is important regardless. But the tradeoff of computing time/cost with problem complexity is hugely important to think about. Finding a proof in a formal language is trivially solvable in theory since you just have to search through possible proofs until you find one ending with the desired statement. The whole practical question is how much time it takes. Three days per problem is, by many standards, a 'reasonable' amount of time. However there are still unanswered questions, notably that 'three days' is not really meaningful in and of itself. How parallelized was the computation; what was the hardware capacity? And how optimized is AlphaProof for IMO-type problems (problems which, among other things, all have short solutions using elementary tools)? These are standard kinds of critical questions to ask.
- dash2 2y agoThough, if you start solving problems that humans can't or haven't solved, then questions of capacity won't matter much. A speedup in the movement of the maths frontier would be worth many power stations.
- nybsjytm 2y agoWhy have they still not released a paper aside from a press release? I have to admit I still don't know how auspicious it is that running google hardware for three days apiece was able to find half-page long solutions, given that the promise has always been to solve the Riemann hypothesis with the click of a button. But of course I do recognize that it's a big achievement relative to previous work in automatic theorem proving.
- whatshisface 2y agoI don't know why so few people realize this, but by solving any of the problems their performance is superhuman for most reasonable definitions of human. Talking about things like solving the Reimman hypothesis in so many years assumes a little too much about the difficulty of problems that we can't even begin to conceive of a solution for. A better question is what can happen when everybody has access to above average reasoning. Our society is structured around avoiding confronting people with difficult questions, except when they are intended to get the answer wrong.
- GregarianChild 2y agoWe know that any theorem that is provable at all (in the chosen foundation of mathematics) can be found by patiently enumerating all possible proofs. So, in order to evaluate AlphaProof's achievements, we'd need to know how much of a shortcut AlphaProof achieved. A good proxy for that would be the total energy usage for training and running AlphaProof. A moderate proxy for that would be the number of GPUs / TPUs that were run for 3 days. If it's somebody's laptop, it would be super impressive. If it's 1000s of TPUs, then less so.
- Onavo 2y ago> We know that any theorem that is provable at all (in the chosen foundation of mathematics) can be found by patiently enumerating all possible proofs. Which computer science theorem is this from?
- deleted 2y ago
- throwaway713 2y agoAnyone else feel like mathematics is sort of the endgame? I.e., once ML can do it better than humans, that’s basically it?
- abrookewood 2y agoI mean ... calculators can do better at mathematics than most of us. I don't think they are going to threaten us anytime soon.
- margorczynski 2y agoI doubt it. Math has the property that you have a way to 100% verify that what you're doing is correct with little cost (as it is done with Lean). Most problems don't have anything close to that.
- exe34 2y agoto be fair, humans also have to run experiments to discover whether their models fit nature - AI will do it too.
- margorczynski 2y agoThese kind of experiments are many times orders of magnitude more costly (time, energy, money, safety, etc.) than verifying a mathematical proof with something like Lean. That's why many think math will be one of the first to crack with AI as there is a relatively cheap and fast feedback loop available.
- AlotOfReading 2y agoMath doesn't have a property that you can verify everything you're doing is correct with little cost. Humans simply tend to prefer theorems and proofs that are simpler.
- thrance 2y agoYou can, in principle, formalize any correct mathematical proof and verify its validity procedurally with a "simple" algorithm, that actually exists (See Coq, Lean...). Coming up with the proof is much harder, and deciding what to attempt to prove even harder, though.
- sega_sai 2y agoI think the interface of LLM with formalized languages is really the future. Because here you can formally verify every statement and deal with hallucinations.
- raincole 2y agoIt's obviously not the future (outside of mathematics research). The whole LLM boom we've seen in the past two years comes from one single fact: peopel don't need to learn a new language to use it.
- seizethecheese 2y agoBoth comments can be right. People don’t need to know HTML to use the internet.
- nickpsecurity 2y agoNatural language -> Formal Language with LLM-assisted tactics/functions -> traditional tools (eg provers/planners) -> expert-readable outputs -> layperson-readable results. I can imagine many uses for flows where LLM’s can implement the outer layers above.
- Groxx 2y agoThe difficulty then will be figuring out if the proof is relevant to what you want, or simply a proof of 1=1 in disguise.
- est 2y ago> formalized languages is really the future Hmm, maybe it's time for symbolism to shine?
- samweb3 2y agoI am building Memelang (memelang.net) to help with this as well. I'd love your thoughts if you have a moment!
- thesz 2y ago
- chompychop 2y agoIs it currently possible to reliably limit the cut-off knowledge of an LLM (either during training or inference)? An interesting experiment would be to feed an LLM mathematical knowledge only up to the year of proving a theorem, and then see if it can actually come up with the novel techniques used in the proof. For example, having only access to papers prior to 1993, can an LLM come up with Wiles' proof of FLT?
- ogrisel 2y agoThat should be doable, e.g. by semi-automated curation of the pre-training dataset. However, since curating such large datasets and running pre-training runs is so expensive, I doubt that anybody will run such an experiment. Especially since would have to trust that the curation process was correct enough for the end-result to be meaningful. Checking that the curation process is not flawed is probably as expensive as running it in the first place.
- n4r9 2y agoThere's the Frontier Math benchmarks [0] demonstrating that AI is currently quite far from human performance at research-level mathematics. [0] https://arxiv.org/abs/2411.04872 https://arxiv.org/abs/2411.04872
- data_maan 2y agoThey didn't demonstrate anything. They haven't even released their dataset, nor mentioned how big it is. It's just hot air, just like the AlphaProof announcement, where very little is know about their system.
- n4r9 2y agoThey won't publish the problem set for obvious reasons. And I doubt it's hot air, given the mathematicians involved in creating it.
- chvid 2y agoMathematicians have been using computers, programming languages, and proof engines for over half a century; however breakthroughs in mathematics are still made by humans in any meaningful sense, even though the tools they use and make are increasingly complex. But as things look now, I will be willing to bet that the next major breakthrough in maths will be touted as being AI/LLMs and coming out of one of the big US tech companies rather than some German university. Why? Simply, the money is much bigger. Such an event would pop the market value of the company involved by a hundred billion - plenty of incentive right there to paint whatever as AI and hire whoever.
- kzrdude 2y agoBut, these AI solutions are trying to solve math problems to prove their AI capabilities, not because they care about mathematics.
- staunton 2y agoSure. Why do you say "but"? Solving such a math problem (while perhaps massively overstating the role AI actually played in the solution) would be great PR for everyone involved.