20 ms·
Can AI do maths yet? Thoughts from a mathematician
- noFaceDiscoG668 2y ago"once" the training data can do it, LLMs will be able to do it. and AI will be able to do math once it comes to check out the lights of our day and night. until then it'll probably wonder continuously and contiguously: "wtf! permanence! why?! how?! by my guts, it actually fucking works! why?! how?!"
- tossandthrow 2y agoI do think it is time to start questioning whether the utility of ai solely can be reduced to the quality of the training data. This might be a dogma that needs to die.
- noFaceDiscoG668 2y agoI tried. I don't have the time to formulate and scrutinise adequate arguments, though. Do you? Anything anywhere you could point me to? The algorithms live entirely off the training data. They consistently fail to "abduct" (inference) beyond any language-in/of-the-training-specific information.
- jstanley 2y agoThe best way to predict the next word is to accurately model the underlying system that is being described.
- tossandthrow 2y agoIt is a gradual thing. Presumably the models are inferring things on runtime that was not a part of their training data. Anyhow, philosophically speaking you are also only exposed to what your senses pick up, but presumably you are able to infer things? As written: this is a dogma that stems from a limited understanding of what algorithmic processes are and the insistence that emergence can not happen from algorithmic systems.
- noFaceDiscoG668 2y ago[dead]
- noFaceDiscoG668 2y ago[dead]
- croes 2y agoIf not bad training data shouldn’t be problem
- kergonath 2y agoThere can be more than one problem. The history of computing (or even just the history of AI) is full of things that worked better and better right until they hit a wall. We get diminishing returns adding more and more training data. It’s really not hard to imagine a series of breakthroughs bringing us way ahead of LLMs.
- Flenkno 2y agoAWS announced 2 or 3 weeks a way of formulating rules into a formal language. AI doesn't need to learn everything, our LLM Models already contain EVERYTHING. Including ways of how to find a solution step by step. Which means, you can tell an LLM to translate whatever you want, into a logical language and use an external logic verifier. The only thing a LLM or AI needs to 'understand' at this point is to make sure that the statistical translation from left to right is high enough. Your brain doesn't just do logic out of the box, You conclude things and formulate them. And plenty of companies work on this. Its the same with programming, if you are able to write code and execute it, you execute it until the compiler errors are gone. Now your LLM can write valid code out of the box. Let the LLM write unit tests, now it can verify itself. Claude for example offers you, out of the box, to write a validation script. You can give claude back the output of the script claude suggested to you. Don't underestimate LLMs
- casenmgreen 2y agoI may be wrong, but I think it a silly question. AI is basically auto-complete. It can do math to the extent you can find a solution via auto-complete based on an existing corpus of text.
- Bootvis 2y agoYou're underestimating the emergent behaviour of these LLM's. See for example what Terrence Tao thinks about o1: https://mathstodon.xyz/@tao/113132502735585408 https://mathstodon.xyz/@tao/113132502735585408
- WhyOhWhyQ 2y agoI'm always just so pleased that the most famous mathematician alive today is also an extremely kind human being. That has often not been the case.
- roflc0ptic 2y agoPretty sure this is out of date now
- noFaceDiscoG668 2y ago[flagged]
- kergonath 2y agoWhy would others provide proofs when you are yourself posting groundless opinions as facts in this very thread?
- noFaceDiscoG668 2y ago[dead]
- mdp2021 2y ago> AI is basically Very many things conventionally labelled in the 50's. You are speaking of LLMs.
- aithrowawaycomm 2y agoI am fairly optimistic about LLMs as a human math -> theorem-prover translator, and as a fan of Idris I am glad that the AI community is investing in Lean. As the author shows, the answer to "Can AI be useful for automated mathematical work?" is clearly "yes." But I am confident the answer to the question in the headline is "no, not for several decades." It's not just the underwhelming benchmark results discussed in the post, or the general concern about hard undergraduate math using different skillsets than ordinary research math. IMO the deeper problem still seems to be a basic gap where LLMs can seemingly do formal math at the level of a smart graduate student but fail at quantitative/geometric reasoning problems designed for fish. I suspect this holds for O3, based on one of the ARC problems it wasn't able to solve: https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90759a16-3602-4aa8-ba8c-69f1d67c31f1_1147x638.webp https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_pr... (via https://www.interconnects.ai/p/openais-o3-the-2024-finale-of-ai https://www.interconnects.ai/p/openais-o3-the-2024-finale-of...) ANNs are simply not able to form abstractions, they can only imitate them via enormous amounts of data and compute. I would say there has been zero progress on "common sense" math in computers since the invention of Lisp: we are still faking it with expert systems, even if LLM expert systems are easier to build at scale with raw data. It is the same old problem where an ANN can attain superhuman performance on level 1 of Breakout, but it has to be retrained for level 2. I am not convinced it makes sense to say AI can do math if AI doesn't understand what "four" means with the same depth as a rat, even if it can solve sophisticated modular arithmetic problems. In human terms, does it make sense to say a straightedge-and-compass AI understands Euclidean geometry if it's not capable of understanding the physical intuition behind Euclid's axioms? It makes more sense to say it's a brainless tool that helps with the tedium and drudgery of actually proving things in mathematics.
- est 2y agoAt this stage I assume everything having a sequencial pattern can and will be automated by LLM AIs.
- Someone 2y agoI think that’s provably incorrect for the current approach to LLMs. They all have a horizon over which they correlate tokens in the input stream. So, for any LLM, if you intersperse more than that number of ‘X’ tokens between each useful token, they won’t be able to do anything resembling intelligence. The current LLMs are a bit like n-gram databases that do not use letters, but larger units.
- red75prime 2y agoThe follow-up question is "Does it require a paradigm shift to solve it?". And the answer could be "No". Episodic memory, hierarchical learnable tokenization, online learning or whatever works well on GPUs.
- beng-nl 2y agoIt’s that a bit of an unfair sabotage? Naturally, humans couldn’t do it, even though they could edit the input to remove the X’s, but shouldn’t we evaluate the ability (even intelligent ability) of LLM’s on what they can generally do rather than amplify their weakness?
- Someone 2y agoWhy is that unfair in reply to the claim “At this stage I assume everything having a sequencial pattern can and will be automated by LLM AIs.”? I am not claiming LLMs aren’t or cannot be intelligent, not even that they cannot do magical things; I just rebuked a statement about the lack of limits of LLMs. > Naturally, humans couldn’t do it, even though they could edit the input to remove the X’s So, what are you claiming: that they cannot or that they can? I think most people can and many would. Confronted with a file containing millions of X’s, many humans will wonder whether there’s something else than X’s in the file, do a ‘replace all’, discover the question hidden in that sea of X’s, and answer it. There even are simple files where most humans would easily spot things without having to think of removing those X's. Consider a file How X X X X X X many X X X X X X days X X X X X X are X X X X X X there X X X X X X in X X X X X X a X X X X X X week? X X X X X X with a million X’s on the end of each line. Spotting the question in that is easy for humans, but impossible for the current bunch of LLMs
- ned99 2y agoI think this is a silly question, you could track AI's doing very simple maths back in 1960 - 1970's
- mdp2021 2y agoIt's just the worrisome linguistic confusion between AI and LLMs.
- jampekka 2y agoI just spent a few days trying to figure out some linear algebra with the help of ChatGPT. It's very useful for finding conceptual information from literature (which for a not-professional-mathematician at least can be really hard to find and decipher). But in the actual math it constantly makes very silly errors. E.g. indexing a vector beyond its dimension, trying to do matrix decomposition for scalars and insisting on multiplying matrices with mismatching dimensions. O1 is a lot better at spotting its errors than 4o but it too still makes a lot of really stupid mistakes. It seems to be quite far from producing results itself consistently without at least a somewhat clueful human doing hand-holding.
- glimshe 2y agoIsn't Wolfram Alpha a better "ChatGPT of Math"?
- Filligree 2y agoWolfram Alpha is better at actually doing math, but far worse at explaining what it’s doing, and why.
- dartos 2y agoWhat’s worse about it? It never tells you the wrong thing, at the very least.
- fn-mote 2y agoIts understanding of problems was very bad last time I used it. Meaning it was difficult to communicate what you wanted it to do. Usually I try to write in the Mathematica language, but even that is not foolproof. Hopefully they have incorporated more modern LLM since then, but it hasn’t been that long.
- jampekka 2y agoWolfram Alpha's "smartness" is often Clippy level enraging. E.g. it makes assumptions of symbols based on their names (e.g. a is assumed to be a constant, derivatives are taken w.r.t. x). Even with Mathematica syntax it tends to make such assumptions and refuses to lift them even when explicitly directed. Quite often one has to change the variable symbols used to try to make Alpha to do what's meant.
- lproven 2y agoBetteridge's Law applies.
- LittleTimothy 2y agoIt's fascinating that this has run into the exact same problem as the Quantum research. Ie, in the quantum research to demonstrate any valuable forward progress you must compute something that is impossible to do with a traditional computer. If you can't do it with a traditional computer, it suddenly becomes difficult to verify correctness (ie, you can't just check it was matching the traditional computer's answer. In the same way ChatGPT scores 25% on this and the question is "How close were those 25% to questions in the training set". Or to put it another way we want to answer the question "Is ChatGPT getting better at applying it's reasoning to out-of-set problems or is it pulling more data into it's training set". Or "Is the test leaking into the training". Maybe the whole question is academic and it doesn't matter, we solve the entire problem by pulling all human knowledge into the training set and that's a massive benefit. But maybe it implies a limit to how far it can push human knowledge forward.
- lazide 2y agoIf constrained by existing human knowledge to come up with an answer, won’t it fundamentally be unable to push human knowledge forward?
- actionfromafar 2y agoThen much of human research and development is also fundamentally impossible.
- AnerealDew 2y agoOnly if you think current "AI" is on the same level as human creativity and intelligence, which it clearly is not.
- actionfromafar 2y agoI think current "AI" (i.e. LLMs) is unable to push human knowledge forward, but not because it's constrained by existing human knowledge. It's more like peeking into a very large magic-8 ball, new answers everytime you shake it. Some useful.
- intellix 2y agoI haven't checked in a while, but last I checked ChatGPT it struggled on very basic things like: how many Fs are in this word? Not sure if they've managed to fix that but since that I had lost hope in getting it to do any sort of math
- sylware 2y agoHow to train an AI strapped to a formal solver.
- puttycat 2y agoNo: https://github.com/0xnurl/gpts-cant-count https://github.com/0xnurl/gpts-cant-count
- sebzim4500 2y agoI can't reliably multiply four digit numbers in my head either, what's your point?
- reshlo 2y agoNobody said you have to do it in your head.
- sebzim4500 2y agoThat's the equivalent to what we are asking the model to do. If you give the model a calculator it will get 100%. If you give it a pen and paper (e.g. let it show it's working) then it will get near 100%.
- reshlo 2y agoCitation needed.
- sebzim4500 2y agoWhich bit do you need a citation for? I can run the experiment in 10 mins.
- reshlo 2y ago> That's the equivalent to what we are asking the model to do. Why? What does it mean to give a model a calculator? What do you mean “let it show its working”? If I ask an LLM to do a calculation, I never said it can’t express the answer to me in long-form text or with intermediate steps. If I ask a human to do a calculation that they can’t reliably do in their head, they are intelligent enough to know that they should use a pen and paper without needing my preemptive permission.
- rishicomplex 2y agoWho is the author?
- williamstein 2y agoKevin Buzzard
- nebulous1 2y agoThere was a little more information in that reddit thread. Of the three difficulty tiers, 25% are T1 (easiest) and 50% are T2. Of the five public problems that the author looked at, two were T1 and two were T2. Glazer on reddit described T1 as "IMO/undergraduate problems", but the article author says that they don't consider them to be undergraduate problems. So the LLM is already doing what the author says they would be surprised about. Also Glazer seemed to regret calling T1 "IMO/undergraduate", and not only because of the disparity between IMO and typical undergraduate. He said that "We bump problems down a tier if we feel the difficulty comes too heavily from applying a major result, even in an advanced field, as a black box, since that makes a problem vulnerable to naive attacks from models" Also, all of the problems shows to Tao were T3
- riku_iki 2y ago> So the LLM is already doing what the author says they would be surprised about. that's if you unconditionally believe in result without any proofreading, confirmation, reproducability and even barely any details (we are given only one slide).
- joe_the_user 2y agoThe reddit thread is ... interesting (direct link[1]). It seems to be a debate among mathematicians some of whom do have access to the secret set. But they're debating publicly and so naturally avoiding any concrete examples that would give the set away so wind-up with fuzzy-fiddly language for the qualities of the problem tiers. The "reality" of keeping this stuff secret 'cause someone would train on it is itself bizarre and certainly shouldn't be above questioning. https://www.reddit.com/r/OpenAI/comments/1hiq4yv/comment/m30yfqp/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button https://www.reddit.com/r/OpenAI/comments/1hiq4yv/comment/m30...
- obastani 2y agoIt's not about training directly on the test set, it's about people discussing questions in the test set online (e.g., in forums), and then this data is swept up into the training set. That's what makes test set contamination so difficult to avoid.
- seafoamteal 2y agoI don't have much to opine from an advanced maths perspective, but I'd like to point out a couple examples of where ChatGPT made basic errors in questions I asked it as an undergrad CS student. 1. I asked it to show me the derivation of a formula for the efficiency of Stop-and-Wait ARQ and it seemed to do it, but a day later, I realised that in one of the steps, it just made a term vanish to get to the next step. Obviously, I should have verified more carefully, but when I asked it to spot the mistake in that step, it did the same thing twice more with bs explanations of how the term is absorbed. 2. I asked it to provide me syllogisms that I could practice proving. An overwhelming number of the syllogisms it gave me were inconsistent and did not hold. This surprised me more because syllogisms are about the most structured arguments you can find, having been formalized centuries ago and discussed extensively since then. In this case, asking it to walk step-by-step actually fixed the issue. Both of these were done on the free plan of ChatGPT, but I can remember if it was 4o or 4.
- voiper1 2y agoThe first question is always: which model? Which fortunately you at least addressed: >free plan of ChatGPT, but I can remember if it was 4o or 4. Since chatgpt-4o, there has been o1-preview, and o1 (full) is out. They just announced o3 got a 25% on frontiermath which is what this article is a reaction to. So, any tests on 4o are at least TWO (or three) AI releases with new capabilities.
- Xcelerate 2y agoSo here's what I'm perplexed about. There are statements in Presburger arithmetic that take time doubly exponential (or worse) in the size of the statement to reach via any path of the formal system whatsoever. These are arithmetic truths about the natural numbers. Can these statements be reached faster in ZFC? Possibly—it's well-known that there exist shorter proofs of true statements in more powerful consistent systems. But the problem then is that one can suppose there are also true short statements in ZFC which likewise require doubly exponential time to reach via any path. Presburger Arithmetic is decidable whereas ZFC is not, so these statements would require the additional axioms of ZFC for shorter proofs, but I think it's safe to assume such statements exist. Now let's suppose an AI model can resolve the truth of these short statements quickly. That means one of three things: 1) The AI model can discover doubly exponential length proof paths within the framework of ZFC. 2) There are certain short statements in the formal language of ZFC that the AI model cannot discover the truth of. 3) The AI model operates outside of ZFC to find the truth of statements in the framework of some other, potentially unknown formal system (and for arithmetical statements, the system must necessarily be sound). How likely are each of these outcomes? 1) is not possible within any coherent, human-scale timeframe. 2) IMO is the most likely outcome, but then this means there are some really interesting things in mathematics that AI cannot discover. Perhaps the same set of things that humans find interesting. Once we have exhausted the theorems with short proofs in ZFC, there will still be an infinite number of short and interesting statements that we cannot resolve. 3) This would be the most bizarre outcome of all. If AI operates in a consistent way outside the framework of ZFC, then that would be equivalent to solving the halting problem for certain (infinite) sets of Turing machine configurations that ZFC cannot solve. That in itself itself isn't too strange (e.g., it might turn out that ZFC lacks an axiom necessary to prove something as simple as the Collatz conjecture), but what would be strange is that it could find these new formal systems efficiently. In other words, it would have discovered an algorithmic way to procure new axioms that lead to efficient proofs of true arithmetic statements. One could also view that as an efficient algorithm for computing BB(n), which obviously we think isn't possible. See Levin's papers on the feasibility of extending PA in a way that leads to quickly discovering more of the halting sequence.
- aleph_minus_one 2y ago
- bambax 2y ago> As an academic mathematician who spent their entire life collaborating openly on research problems and sharing my ideas with other people, it frustrates me [that] I am not even to give you a coherent description of some basic facts about this dataset, for example, its size. However there is a good reason for the secrecy. Language models train on large databases of knowledge, so you moment you make a database of maths questions public, the language models will train on it. Well, yes and no. This is only true because we are talking about closed models from closed companies like so-called "OpenAI". But if all models were truly open, then we could simply verify what they had been trained on, and make experiments with models that we could be sure had never seen the dataset. Decades ago Microsoft (in the words of Ballmer and Gates) famously accused open source of being a "cancer" because of the cascading nature of the GPL. But it's the opposite. In software, and in knowledge in general, the true disease is secrecy.
- ludwik 2y ago> But if all models were truly open, then we could simply verify what they had been trained on How do you verify what a particular open model was trained on if you haven’t trained it yourself? Typically, for open models, you only get the architecture and the trained weights. How can you reliably verify what the model was trained on from this? Even if they provide the training set (which is not typically the case), you still have to take their word for it—that’s not really "verification."
- asadotzler 2y agoThe OP said "truly open" not "open model" or any of the other BS out there. If you are truly open you share the training corpora as well or at least a comprehensive description of what it is and where to get it.
- ludwik 2y agoIt seems like you skipped the second paragraph of my comment?
- 4ad 2y ago> FrontierMath is a secret dataset of “hundreds” of hard maths questions, curated by Epoch AI, and announced last month. The database stopped being secret when it was fed to proprietary LLMs running in the cloud. If anyone is not thinking that OpenAI has trained and tuned O3 on the "secret" problems people fed to GPT-4o, I have a bridge to sell you.
- fn-mote 2y agoThis level of conspiracy thinking requires evidence to be useful. Edit: I do see from your profile that you are a real person though, so I say this with more respect.
- dns_snek 2y agoWhat evidence do we need that AI companies are exploiting every bit of information they can use to get ahead in the benchmarks to generate more hype? Ignoring terms/agreements, violating copyright, and otherwise exploiting information for personal gain is the foundation of that entire industry for crying out loud.
- threeseed 2y agoSome people are also forgetting who is the CEO of OpenAI. Sam Altman has long talked about believing in the "move fast and break things" way of doing business. Which is just a nicer way of saying do whatever dodgy things you can get away with.
- cheald 2y agoOpenAI's also in the position of having to compete against other LLM trainers - including the open-weights Llama models and their community derivatives, which have been able to do extremely well with a tiny fraction of OpenAI's resources - and to justify their astronomical valuation. The economic incentive to cheat is extreme; I think that cheating has to be the default presumption.
- advisedwang 2y ago
- ashoeafoot 2y agoAi has a interior world model thus it can do math if a chain of proof is walking without uncertainty from room to room. the problem is its inability to reflect on its own uncertainty and to then overrife that uncertainty ,should a new room entrance method be selfsimilar to a previous entrance
- voidhorse 2y agoEventually we may produce a collection of problems exhaustive enough that these tools can solve almost any problem that isn't novel in practice, but I doubt that they will ever become general problem solvers capable of what we consider to be reasoning in humans. Historically, the claim that neural nets were actual models of the human brain and human thinking was always epistemically dubious. It still is. Even as the practical problems of producing better and better algorithms, architectures, and output have been solved, there is no reason to believe a connection between the mechanical model and what happens in organisms has been established. The most important point, in my view, is that all of the representation and interpretation still has to happen outside the computational units. Without human interpreters, none of the AI outputs have any meaning. Unless you believe in determinism and an overseeing god, the story for human beings is much different. AI will not be capable of reason until, like humans, it can develop socio-rational collectivities of meaning that are independent of the human being. Researchers seemed to have a decent grasp on this in the 90s, but today, everyone seems all too ready to make the same ridiculous leaps as the original creators of neural nets. They did not show, as they claimed, that thinking is reducible to computation. All they showed was that a neural net can realize a boolean function—which is not even logic, since, again, the entire semantic interpretive side of the logic is ignored.
- nmca 2y agoCan you define what you mean by novel here?
- red75prime 2y ago> there is no reason to believe a connection between the mechanical model and what happens in organisms has been established The universal approximation theorem. And that's basically it. The rest is empirical. No matter which physical processes happen inside the human brain, a sufficiently large neural network can approximate them. Barring unknowns like super-Turing computational processes in the brain.
- lupire 2y agoThat's not useful by itself, because "anything cam model anything else" doesn't put any upper bound on emulation cost, which for one small task could be larger than the total energy available in the entire Universem
- alphan0n 2y agoAs far as ChatGPT goes, you may as well be asking: Can AI use a calculator? The answer is yes, it can utilize a stateful python environment and solve complex mathematical equations with ease.
- lcnPylGDnU4H9OF 2y agoThere is a difference between correctly stating that 2 + 2 = 4 within a set of logical rules and proving that 2 + 2 = 4 must be true given the rules.
- alphan0n 2y agoI think you misunderstood, ChatGPT can utilize Python to solve a mathematical equation and provide proof. https://chatgpt.com/share/676980cb-d77c-8011-b469-4853647f9804 https://chatgpt.com/share/676980cb-d77c-8011-b469-4853647f98... More advanced solutions: https://chatgpt.com/share/6769895d-7ef8-8011-8171-6e84f3310366 https://chatgpt.com/share/6769895d-7ef8-8011-8171-6e84f33103...
- cruffle_duffle 2y agoIt still has to know what to code in that environment. And based on my years of math as a wee little undergrad, the actual arithmetic was the least interesting part. LLM’s are horrible at basic arithmetic, but they can use python for the calculator. But python wont help them write the correct equations or even solve for the right thing (wolfram alpha can do a bit of that though)
- alphan0n 2y agoYou’ll have to show me what you mean. I’ve yet to encounter an equation that 4o couldn’t answer in 1-2 prompts unless it timed out. Even then it can provide the solution in a Jupyter notebook that can be run locally.
- cruffle_duffle 2y agoNever really pushed it. I have to reason to believe it wouldn’t get most of that stuff correctly. Math is very much like programming and I’m sure it can output really good python for its notebook to use execute.
- upghost 2y agoI didn't see anyone else ask this but.. isn't the FrontierMath dataset compromised now? At the very least OpenAI now knows the questions if not the answers. I would expect that the next iteration will "magically" get over 80% on the FrontierMath test. I imagine that experiment was pretty closely monitored.
- jvanderbot 2y agoI figured their model was independently evaluated against the questions/answers. That's not to say it's not compromised by "Here's a bag of money" type methods, but I don't even think it'd be a reasonable test if they just handed over the dataset.
- upghost 2y agoI'm sure it was independently evaluated, but I'm sure the folks running the test were not given an on-prem installation of ChatGPT to mess with. It was still done via API calls, presumably through the chat interface UI. That means the questions went over the fence to OpenAI. I'm quite certain they are aware of that, and it would be pretty foolish not to take advantage of at least knowing what the questions are.
- jvanderbot 2y agoNow that you put it that way, it is laughably easy.
- ls612 2y agoDepending on the plan the researchers used they may have contractual protections against OpenAI training on their inputs.
- upghost 2y agoSure, but given the resourcing at OpenAI, it would not be hard to clean[1] the inputs. I'm just trying to be realistic here, there are plenty of ways around contractual obligations and a significant incentive to do so. [1]: https://en.wikipedia.org/wiki/Clean-room_design https://en.wikipedia.org/wiki/Clean-room_design
- sincerecook 2y agoNo it can't, and there's no such thing as AI. How is a thing that predicts the next-most-likely word going to do novel math? It can't even do existing math reliably because logical operations and statistical approximation are fundamentally different. It is fun watching grifters put lipstick on this thing and shop it around as a magic pig though.
- bwfan123 2y agoopenai and epochai (frontier math) are startups with a strong incentive to push such narratives. the real test will be in actual adoption in real world use cases. the management class has a strong incentive to believe in this narrative, since it helps them reduce labor cost. so they are investing in it. eventually, the emperor will be seen to have no clothes at least in some usecases for which it is being peddled right now.
- comp_throw7 2y agoEpoch is a non-profit research institute, not a startup.
- retrocryptid 2y agoWhen did we decide that AI == LLM? Oh don't answer. I know, The VC world noticed CNNs and LLMs about 10 years ago and it's the only thing anyone's talked about ever since. Seems to me the answer to 'Can AI do maths yet?' depends on what you call AI and what you call maths. Our old departmental VAX running at a handfull of megahertz could do some very clever symbol manipulation on binomials and if you gave it a few seconds, it could even do something like theorum proving via proto-prolog. Neither are anywhere close to the glorious GAI future we hope to sell to industry and government, but it seems worth considering how they're different, why they worked, and whether there's room for some hybrid approach. Do LLMs need to know how to do math if they know how to write Prolog or Coc statements that can do interesting things? I've heard people say they want to build software that emulates (simulates?) how humans do arithmetic, but ask a human to add anything bigger than two digit numbers and the first thing they do is reach for a calculator.
- aaron695 2y ago[dead]
- caroline0v0 2y agoIn fact, what I am most curious about is how AI understands symbolic logic relationships (neural networks and Turing machines are not completely equivalent). During training, this is a bunch of tokens.
- skydhash 2y agoI wouldn't say understand. But your answers is patterns. Formalism is mostly definition (axioms) and inference rules (theories). If we take programming languages, most grammars (which describe these two elements) are only a few pages long. With LLM being patterns seeker at its core, I guess it would be easy to extract the rules from a sample of programs, as the structure is so rigid. You won't get the Turing machine evaluation mechanism and determinism, but you will have a generator. Although the viability of what is generated is is question. Because the other part of formalism, semantics, is almost always missing.
- ivansavz 2y agoYesterday, I saw a thought provoking talk about the future of of "math jobs" assuming automated theory proving becomes more prevalent in the future. [ (Re)imagining mathematics in a world of reasoning machines by Akshay Venkatesh] https://www.youtube.com/watch?v=vYCT7cw0ycw https://www.youtube.com/watch?v=vYCT7cw0ycw [54min] Abstract: In the coming decades, developments in automated reasoning will likely transform the way that research mathematics is conceptualized and carried out. I will discuss some ways we might think about this. The talk will not be about current or potential abilities of computers to do mathematics—rather I will look at topics such as the history of automation and mathematics, and related philosophical questions. See discussion at https://news.ycombinator.com/item?id=42465907 https://news.ycombinator.com/item?id=42465907
- qnleigh 2y agoThat was wonderful, thank you for linking it. For the benefit of anyone who doesn't have time to watch the whole thing, here are a few really nice quotes that convey some main points. "We might put the axioms into a reasoning apparatus like the logical machinery of Stanley Jevons, and see all geometry come out of it. That process of reasoning are replaced by symbols and formulas... may seem artificial and puerile; and it is needless to point out how disastrous it would be in teaching and how hurtful to the mental development; how deadening it would be for investigators, whose originality it would nip in the bud. But as used by Professor Hilbert, it explains and justifies itself if one remembers the end pursued." Poincare on the value of reasoning machines, but the analogy to mathematics once we have theorem-proving AI is clear (that the tools and the lie direct outputs are not the ends. Human understanding is). "Even if such a machine produced largely incomprehensible proofs, I would imagine that we would place much less value on proofs as a goal of math. I don't think humans will stop doing mathematics... I'm not saying there will be jobs for them, but I don't think we'll stop doing math." "Mathematics is the study of reproducible mental objects." This definition is human ("mental") and social (it implies reproducing among individuals). "Maybe in this world, mathematics would involve a broader range of inquiry... We need to renegotiate the basic goals and values of the discipline." And he gives some examples of deep questions we may tackle beyond just proving theorems.
- swalsh 2y agoEvery profession seems to have a pessimistic view of AI as soon as it starts to make progress in their domain. Denial, Anger, Bargaining, Depression, and Acceptance. Artists seem to be in the depression state, many programmers are still in the denial phase. Pretty solid denial here from a mathematician. o3 was a proof of concept, like every other domain AI enters, it's going to keep getting better. Society is CLEARLY not ready for what AI's impact is going to be. We've been through change before, but never at this scale and speed. I think Musk/Vivek's DOGE thing is important, our governent has gotten quite large and bureaucratic. But the clock has started on AI, and this is a social structural issue we've gotta figure out. Putting it off means we probably become subjects to a default set of rulers if not the shoggoth itself.
- haolez 2y agoI think it's a little of both. Maybe generative AI algorithms won't overcome their initial limitations. But maybe we don't need to overcome them to transform society in a very significant way.
- WanderPanda 2y agoOr is it just white collar workers experiencing what blue collar workers have been experiencing for decades?
- mensetmanusman 2y agoThe reason why this is so disruptive is because it will effect hundreds of fields simultaneously. Previously workers in a field disrupted by automation would retrain to a different part of the economy. If AI pans out to the point that there are mass layoffs in hundreds of sectors of the economy at once, then i’m not sure the process we have haphazardly set up now will work. People will have no idea where to go beyond manual labor. (But this will be difficult due to the obesity crisis - but maybe it will save lives in a weird way).
- jebarker 2y ago> I am dreading the inevitable onslaught in a year or two of language model “proofs” of the Riemann hypothesis which will just contain claims which are vague or inaccurate in the middle of 10 pages of correct mathematics which the human will have to wade through to find the line which doesn’t hold up. I wonder what the response of working mathematicians will be to this. If the proofs look credible it might be too tempting to try and validate them, but if there’s a deluge that could be a hug time sync. Imagine if Wiles or Perelman had produced a thousand different proofs for their respective problems.
- bqmjjx0kac 2y agoMaybe the coming onslaught of AI slop "proofs" will give a little bump to proof assistants like Coq. Of course, it would still take a human mathematician some time to verify theorem definitions.
- Hizonner 2y agoDon't waste time on looking at it unless a formal proof checker can verify it.
- kevinventullo 2y agoHonestly I think it won’t be that different from today, where there is no shortage of cranks producing “proofs” of the Riemann Hypothesis and submitting them to prestigious journals.
- yodsanklai 2y agoI understand the appeal of having a machine helping us with maths and expanding the frontier of knowledge. They can assist researchers and make them more productive. Just like they can make already programmers more productive. But maths are also fun and fulfilling activity. Very often, when we learn a math theory, it's because we want to understand and gain intuition on the concepts, or we want to solve a puzzle (for which we can already look up the solution). Maybe it's similar to chess. We didn't develop search engines to replace human players and make them play together, but they helped us become better chess players or understanding the game better. So the recent progress is impressive, but I still don't see how we'll use this tech practically and what impacts it can have and in which fields.
- vouaobrasil 2y agoMy favourite moments of being a graduate student in math was showing my friends (and sometimes professors) proofs of propositions and theorems that we discussed together. To be the first to put together a coherent piece of reasoning that would convince them of the truth was immensely exciting. Those were great bonding moments amongst colleagues. The very fact that we needed each other to figure out the basics of the subject was part of what made the journey so great. Now, all of that will be done by AI. Reminds of the time when I finally enabled invincibility in Goldeneye 007. Rather boring. I think we've stopped to appreciate the human struggle and experience and have placed all the value on the end product, and that's we're developing AI so much. Yeah, there is the possibility of working with an AI but at that point, what is the point? Seems rather pointless to me in an art like mathematics.
- sourcepluck 2y ago> Now, all of that will be done by AI. No "AI" of any description is doing novel proofs at the moment. Not o3, or anything else. LLMs are good for chatting about basic intuition with, up to and including complex subjects, if and only if there are publically available data on the topic which have been fed to the LLM during its training. They're good at doing summaries and overviews of specific things (if you push them around and insist they don't waffle and ignore garbage carefully and keep your critical thinking hat on, etc etc). It's like having a magnifying glass that focuses in on the small little maths question you might have, without you having to sift through ten blogs or videos or whatever. That's hardly going to replace graduate students doing proofs with professors, though, at least not with the methods being employed thus far!
- vouaobrasil 2y agoI am talking about in 20-30 years.
- busyant 2y agoAs someone who has a 18 yo son who wants to study math, this has me (and him) ... worried ... about becoming obsolete? But I'm wondering what other people think of this analogy. I used to be a bench scientist (molecular genetics). There were world class researchers who were more creative than I was. I even had a Nobel Laureate once tell me that my research was simply "dotting 'i's and crossing 't's". Nevertheless, I still moved the field forward in my own small ways. I still did respectable work. So, will these LLMs make us completely obsolete? Or will there still be room for those of us who can dot the "i"?--if only for the fact that LLMs don't have infinite time/resources to solve "everything." I don't know. Maybe I'm whistling past the graveyard.
- deepsun 2y agoBy the way, don't trust Nobel laureates or even winners. E.g. Linus Pauling was talking absolute garbage, harmful and evil, after winning the Nobel.
- Radim 2y ago> don't trust Nobel laureates or even winners Nobel laureate and winner are the same thing. > Linus Pauling was talking absolute garbage, harmful and evil, after winning the Nobel. Can you be more specific, what garbage? And which Nobel prize do you mean – Pauling got two, one for chemistry and one for peace.
- bongodongobob 2y agoEugenics and vitamin C as a cure all.
- lern_too_spel 2y agoIf Pauling's eugenics policies were bad, then the laws against incest that are currently on the books in many states (which are also eugenics policies that use the same mechanism) are also bad. There are different forms of eugenics policies, and Pauling's proposal to restrict the mating choices of people carrying certain recessive genes so their children don't suffer is ethically different from Hitler exterminating people with certain genes and also ethically different from other governments sterilizing people with certain genes. He later supported voluntary abortion with genetic testing, which is now standard practice in the US today, though no longer in a few states with ethically questionable laws restricting abortion. This again is ethically different from forced abortion. https://scarc.library.oregonstate.edu/coll/pauling/blood/narrative/page35.html https://scarc.library.oregonstate.edu/coll/pauling/blood/nar...
- jokoon 2y agoI wish scientists who do psychology and cognition of actual brains could approach those AI things and talk about it, and maybe make suggestions. I really really wish AI would make some breakthrough and be really useful, but I am so skeptical and negative about it.
- joe_the_user 2y agoUnfortunately, the scientists who study actually brains have all sort of interesting models but ultimately very little clue how these actual brains work at the level of problem solving. I mean, there's all sort of "this area is associated with that kind of process" and "here's evidence this area does this algorithm" stuff but it's all at the level you imagine steam engine engineers trying to understand a warp drive. The "open worm project" was an effort years ago to get computer scientists involved in trying to understand what "software" a very small actual brain could run. I believe progress here has been very slow and that an idea of ignorance that much larger brains involve. https://en.wikipedia.org/wiki/OpenWorm https://en.wikipedia.org/wiki/OpenWorm
- bongodongobob 2y agoIf you can't find useful things for LLMs or AI at this point, you must just lack imagination.
- 0points 2y ago> How much longer this will go on for nobody knows, but there are lots of people pouring lots of money into this game so it would be a fool who bets on progress slowing down any time soon. Money cannot solve the issues faced by the industry which mainly revolves around lack of training data. They already used the entirety of the internet, all available video, audio and books and they are now dealing with the fact that most content online is now generated by these models, thus making it useless as training data.
- charlieyu1 2y agoOne thing I know is that there wouldn’t be machines entering IMO 2025. The concept of “marker” does not exist in IMO - scores are decided by negotiations between team leaders of each country and the juries. It is important to get each team leader involved for grading the work of students for their country, for accountability as well as acknowledging cultural differences. And the hundreds of people are not going to stay longer to grade AI work.
- witnesser2 2y agoI was not refuted sufficiently a couple of years ago. I claimed "training is open boundary" etc.
- witnesser2 2y agoLike as a few years ago, I just boringly add again "you need modeling" to close it.
- mangomountain 2y agoIn other news we’ve discovered life (our bacteria) on mars Just joking
- Syzygies 2y ago"Can AI do math for us" is the canonical wrong question. People want self-driving cars so they can drink and watch TV. We should crave tools that enhance our abilities, as tools have done since prehistoric times. I'm a research mathematician. In the 1980's I'd ask everyone I knew a question, and flip through the hard bound library volumes of Mathematical Reviews, hoping to recognize something. If I was lucky, I'd get a hit in three weeks. Internet search has shortened this turnaround. One instead needs to guess what someone else might call an idea. "Broken circuits?" Score! Still, time consuming. I went all in on ChatGPT after hearing that Terry Tao had learned the Lean 4 proof assistant in a matter of weeks, relying heavily on AI advice. It's clumsy, but a very fast way to get suggestions. Now, one can hold involved conversations with ChatGPT or Claude, exploring mathematical ideas. AI is often wrong, never knows when it's wrong, but people are like this too. Read how the insurance incidents for self-driving taxis are well below the human incident rates? Talking to fellow mathematicians can be frustrating, and so is talking with AI, but AI conversations go faster and can take place in the middle of the night. I don't want AI to prove theorems for me, those theorems will be as boring as most of the dreck published by humans. I want AI to inspire bursts of creativity in humans.
- ninetyninenine 2y agoYour optimism should be tempered with the downside of progress meaning that AI in the near future may not only inspire creativity in humans, but it can replace human creativity all together. Why do I need to hire an artist for my movie/video game/advertisement when AI can replicate all the creativity I need.
- wnc3141 2y agoThere is research on AI limiting creative output in completive arenas. Essentially it breaks expectancy therefore deteriorates iteration. https://direct.mit.edu/rest/article-abstract/102/3/583/96779/Creativity-Under-Fire-The-Effects-of-Competition?redirectedFrom=fulltext https://direct.mit.edu/rest/article-abstract/102/3/583/96779...
- immibis 2y agoThis was about mathematics.
- Chengdavid 2y ago[dead]
- Onavo 2y agoConsidering that they have Terence Tao himself working on the problem, betting against it would be unwise.
- Sparkyte 2y agoAfter playing with and using AI for almost two years now it is not getting better from both a cost perspective and performance. So the higher the cost the better the performance. While models and hardware can be improved the curve is still steep. The big answer is what are people using it for? We'll they are using lightweight simplistic models to do targeted tasks. To do many smaller and easier to process tasks. Most of the news on AI is just there to promote a product to earn more cash.
- aomix 2y agoNo comment on the article it's just always interesting to get hit with intense jargon from a field I know very little about. I understood the statements of all five questions. I could do the third one relatively quickly (I had seen the trick before that the function mapping a natural n to alpha^n was p-adically continuous in n iff the p-adic valuation of alpha-1 was positive)
- gldjmp 2y agoHaha, the thing about jargon is that is typically hiding something not so bad. (At least in this case, the solution is something you could explain to a high-schooler).
- YeGoblynQueenne 2y ago>> There were language models before ChatGPT, and on the whole they couldn’t even write coherent sentences and paragraphs. ChatGPT was really the first public model which was coherent. If that's referring to Large Language Models, meaning everything after the fist GPT and BERT, then that's absolutely not right. The first LLM that demonstrated the ability to generate coherent, fluently grammatical English was GPT-2. That story about the unicorns- that was the first time a statistical language model was able to generate text that stayed on the subject over a long distance and made (some) sense. GPT-2 was followed by GPT 3 and GPT 3.5 that turned the hype dial up to 11 and were certainly "public" at least if that means publicly available. They were coherent enough that many people predicted all sorts of fancy things, like the end of programming jobs and the end of journalist jobs and so on. So, weird statement that one and it kind of makes me wary of Gell-Mann amnesia while reading the article.
- a_petrov 2y agoI use ChatGPT to help study linear algebra. It helps me a lot when I feel lost. It's often wrong in the calculations, but it's cool to have a study buddy that doesn't judge you. If I get blocked with a problem I can't solve, I ask for assistance with my approach. I enjoy asking ChatGPT about the context behind all that math theory. It's nice to elaborate on that as most of the math books are very lean and provide no applied context.
- vnjxk 2y agoIsn't there a math theorem programming language or something?