11 ms·
I feel very comfortable saying, as a mathematician, that the ability to solve grade school maths problems would not be at all a predictor of ability to solve re
by wbhart 3y ago
I feel very comfortable saying, as a mathematician, that the ability to solve grade school maths problems would not be at all a predictor of ability to solve real mathematical problems at a research level.
The reason LLMs fail at solving mathematical problems is because:
1) they are terrible at arithmetic,
2) they are terrible at algebra, but most importantly,
3) they are terrible at complex reasoning (more specifically they mix up quantifiers and don't really understand the complex logical structure of many arguments)
4) they (current LLMs) cannot backtrack when they find that what they already wrote turned out not to lead to a solution, and it is too expensive to give them the thousands of restarts they'd require to randomly guess their way through the problem if you did give them that facility
Solving grade-school problems might mean progress in 1 and 2, but that is not at all impressive, as there are perfectly good tools out there that solve those problems just fine, and old-style AI researchers have built perfectly good tools for 3. The hard problem to solve is problem 4, and this is something you teach people how to do at a university level.
(I should add that another important problem is what is known as premise selection. I didn't list that because LLMs have actually been shown to manage this ok in about 70% of theorems, which basically matches records set by other machine learning techniques.)
(Real mathematical research also involves what is known as lemma conjecturing. I have never once observed an LLM do it, and I suspect they cannot do so. Basically the parameter set of the LLM dedicated to mathematical reasoning is either large enough to model the entire solution from end to end, or the LLM is likely to completely fail to solve the problem.)
I personally think this entire article is likely complete bunk.
Edit: after reading replies I realise I should have pointed out that humans do not simply backtrack. They learn from failed attempts in ways that LLMs do not seem to. The material they are trained on surely contributes to this problem.
- whatever1 3y agoThe thing is that a LLMs can point out a logic error in their reasoning if specifically asked to do so. So maybe OpenAI just slapped an RL agent on top of the next-token generator.
- abeppu 3y agoThis comment seems to presume that Q* is related to existing LLM work -- which isn't stated in the article. Others have guessed that the 'Q' in Q* is from Q-learning in RL. In particular backtracking, which you point out LLMs cannot do, would not be an issue in an appropriate RL setup.
- nullc 3y ago> which you point out LLMs cannot do, would not be an issue in an appropriate RL setup. Hm? it's pretty trivial to use a sampler for LLMs that has a beam search and will effectively 'backtrack' a 'bad' selection. It just doesn't normally help-- by construction the LLM sampled normally already approximates the correct overall distribution for the entire output, without any search. I assume using a beam search does help when your sampler does have some non-trivial constraints (like the output satisfies some grammar or passes an algebraic test, or even just top-n sampling since those adjustments on a token by token basis result in a different approximate distribution than the original distribution filtered by the constraints).
- nijave 3y agoChatGPT (3.5) seems to do some rudimentary backtracking when told it's wrong enough times. However, it does seem to do very poorly in the logic department. LLMs can't seem to pick out nuance and separate similar ideas that are technically/logically different. They're good at putting things together commonly found together but not so good at separating concepts back out into more detailed sub pieces.
- wbhart 3y agoI've tested GPT-4 on this and it can be induced to give up on certain lines of argument after recognising they aren't leading anywhere and to try something else. But it would require thousands (I'm really under exaggerating here) of restarts to get through even fairly simple problems that professional mathematicians solve routinely. Currently the context length isn't even long enough for it to remember what problem it was solving. And I've tried to come up with a bunch of ways around this. They all fail for one reason or another. LLMs are really a long, long way off managing this efficiently in my opinion.
- Davidzheng 3y agoWeird time estimate given that a little more than a year ago, the leading use of LLMs was generating short coherent paragraphs (3-4 sentences)
- nijave 3y agoI've pasted docs and error messages into GPT 3.5 and it's admitted it's wrong but usually it'll go through a few different answers before returning back to the original and looping
- muskmusk 3y agoFriend, the creator of this new progress is a machine learning PhD with a decade of experience in pushing machine learning forward. He knows a lot of math too. Maybe there is a chance that he too can tell the difference between a meaningless advance and an important one?
- neilk 3y agoI am neither a mathematician or LLM creator but I do know how to evaluate interesting tech claims. The absolute best case scenario for a new technology is that it when it seems like a toy for nerds, and doesn't outperform anything we have today, but the scaling path is clear. Its problems just won't matter if it does that one thing with scaling. The web is a pretty good hypermedia platform, but a disastrously bad platform for most other computer applications. Nevertheless the scaling of URIs and internet protocols have caused us to reorganize our lives around it. And then if there really are unsolvable problems with the platform they just get offloaded onto users. Passwords? Privacy? Your problem now. Surely you know to use a password manager? I think this new wave of AI is going to be like that. If they never solve the hallucination/confabulation issue, it's just going to become your problem. If they never really gain insight, it's going to become your problem to instruct them carefully. Your peers will chide for not using a robust AI-guardrail thing or not learning the basics of prompt engineering like all the kids do instinctively these days.
- wbhart 3y agoHow on earth could you evaluate the scaling path with too little information. That's my point. You can't possibly know that a technology can solve a given kind of problem if it can only so far solve a completely different kind of problem which is largely unrelated! Saying that performance on grade-school problems is predictive of performance on complex reasoning tasks (including theorem proving) is like saying that a new kind of mechanical engine that has 90% efficiency can be scaled 10x. These kind of scaling claims drive investment, I get it. But to someone who understands (and is actually working on) the actual problem that needs solving, this kind of claim is perfectly transparent!
- Davidzheng 3y agoI agree that in and of itself it's not enough to be alarmed. Also i have to say i don't really know what grade school mathematics means here(multiplication? Proving triangles are congruent?). But I think the question is, whether the breakthrough is an algorithmic change in reasoning. If it is, then it could challenge all 4 of your limitations. Again this article is low on details so really we are arguing over our best guesses. But I wouldn't be so confident that an improvement on simple math problems due to algorithms can have huge implications. Also, do you remember what go players said when they beat Fan Hui? Change can come quick
- wbhart 3y agoI think maybe I didn't make myself quite clear here. There are already algorithms which can solve advanced mathematical problems 100% reliably (prove theorems). There are even algorithms which can prove any correct theorem that can be stated in a certain logical language, given enough time. There are even systems in which these algorithms have actually been implemented. My point is that no technology which can solve grade school maths problems would be viewed as a breakthrough by anyone who understood the problem. The fundamental problems which need to be solved are not problems you encounter in grade school mathematics. The article is just ill-informed.
- himaraya 3y agoThe article suggests the way Q* solves basic math problems matters more than the difficulty of the problems themselves. Either way, I think judging the claims made remains premature without seeing the supporting documentation.
- kenjackson 3y ago“Given enough time” makes that a useless statement. Every kid in college learns this. The ability to eventually solve a given theorem isn’t interesting — especially if the time is longer than the time left in the universe. It’s far more interesting to see if an AI can, given an arbitrarily stated problem make clear progress quickly.
- tim333 3y ago
- adastra22 3y agoBack-tracking is a very nearly solved problem in the context of Prolog-like languages or mathematical theorem provers (as you probably well know). There are many ways you could integrate an LLM-like system into a tactic-based theorem prover without having to restart from the beginning for each alternative. Simply checkpointing and backtracking to a checkpoint would naively improve upon your described Monte Carlo algorithm. More likely I assume they are using RL to unwind state backwards and update based on the negative result, which would be significantly more complicated but also much more powerful (essentially it would one-shot learn from each failure). That's just what I came up with after thinking on it for 2 minutes. I'm sure they have even better ideas.
- wbhart 3y agoThere are certainly efforts along the lines of what you suggest. There are problems though. The number of backtracks is 10^k where k is not 2, or 3, or 4..... Another issue is that of autoformalisation. This is the one part of the problem where an LLM might be able to help, if it were reliable enough (it isn't currently) or if it could truly understand the logical structure of mathematical problems correctly (currently they can't).
- visarga 3y agoYou can also consider the chatGPT app as a RL environment. The environment is made of the agent (AI), a second agent (human), and some tools (web search, code, APIs, vision). This grounds the AI into human and tool responses. They can generate feedback that can be incorporated into the model by RL methods. Basically every reply from a human can be interpreted as a reward signal. If the human restates the question, it means a negative reward, the AI didn't get it. If the human corrects the AI, another negative reward, but if they continue the thread then it is positive. You can judge turn-by-turn and end-to-end all chat logs with GPT4 to annotate. The great thing about chat based feedback is that it is scalable. OpenAI has 100M users, they generate these chat sessions by the millions every day. Then they just need to do a second pass (expensive, yes) to annotate the chat logs with RL reward signals and retrain. But they get the human-in-the-loop for free, and that is the best source of feedback. AI-human chat data is in-domain for both the AI and human, something we can't say about other training data. It will contain the kind of mistakes AI does, and the kind of demands humans want to solve with AI. My bet is that OpenAI have realized this and created GPTs in order to enrich and empower the AI to create the best training data for GPT-5. The secret sauce of OpenAI is not their people, or Sam, or the computers, but the training set, especially the augmented and synthetic parts.
- caesil 3y agoFWIW The Verge is reporting that people inside are also saying the Reuters story is bunk: https://www.theverge.com/2023/11/22/23973354/a-recent-openai-breakthrough-on-the-path-to-agi-has-caused-a-stir https://www.theverge.com/2023/11/22/23973354/a-recent-openai...
- himaraya 3y ago> After being contacted by Reuters, OpenAI, which declined to comment, acknowledged in an internal message to staffers a project called Q* and a letter to the board before the weekend's events, one of the people said. Reuters update 6:51 PST The Verge has acted like an intermediary for Sam's camp during this whole saga, from my reading.
- CamperBob2 3y ago"Also, Crysis runs like crap on my Commodore 64."
- calf 3y agoBut, isn't AlphaGo a solution to kind of specific mathematical problem? And that it has passed with flying colors? What I mean is, yes, neural networks are stochastic and that seems to be why they're bad at logic; on the other hand it' not exactly hallucinating a game of Go, and that seems different to how neural networks are prone to hallucination and confabulation on natural language or X-ray imaging.
- wbhart 3y agoSure, but people have already applied deep learning techniques to theorem proving. There are some impressive results (which the press doesn't seem at all interested in because it doesn't have ChatGPT in the title). It's really harder than one might imagine to develop a system which is good at higher order logic, premise selection, backtracking, algebraic manipulation, arithmetic, conjecturing, pattern recognition, visual modeling, has a good mathematical knowledge, is autonomous and fast enough to be useful. For my money, it isn't just a matter of fitting a few existing jigsaw pieces together in some new combination. Some of the pieces don't exist yet.
- visarga 3y agoYou seem knowledgeable. Can you share a couple of interesting papers for theorem proving that came out in the last year? I read a few of them as they came out, and it seemed neural nets can advance the field by mixing "soft" language with "hard" symbolic systems.
- wbhart 3y agoThe field is fairly new to me. I'm originally from computer algebra, and somehow struggling into the field of ATP. The most interesting papers to me personally are the following three: * Making higher order superposition work. https://doi.org/10.1007/978-3-030-79876-5_24 https://doi.org/10.1007/978-3-030-79876-5_24 * MizAR 60 for Mizar 50. https://doi.org/10.48550/arXiv.2303.06686 https://doi.org/10.48550/arXiv.2303.06686 * Magnus Hammer, a Transformer Based Approach to Premise Selection. https://doi.org/10.48550/arXiv.2303.04488 https://doi.org/10.48550/arXiv.2303.04488 Your mileage may vary.
- xcv123 3y ago> I feel very comfortable saying, as a mathematician, that the ability to solve grade school maths problems would not be at all a predictor of ability to solve real mathematical problems at a research level. At some point in the past, you yourself were only capable of solving grade school maths problems.
- SantalBlush 3y agoThe statement you quoted also holds for humans. Of those who can solve grade school math problems, very, very few can solve mathematical problems at a research level.
- xcv123 3y agoYou missed the point. Deep learning models are in the early stages of development. With recent advancements they can already outperform humans at many tasks that were considered to require AGI level machine intelligence just a few years ago.
- deleted 3y ago[deleted]
- kgeist 3y agoWe're moving the goalposts all the time. First we had the Turing test, now AI solving math problems "isn't impressive". Any small mistake is a proof it cannot reason at all. Meanwhile 25% humans think the Sun revolves around the Earth and 50% of students get the bat and ball problem wrong.
- deleted 3y ago[deleted]
- cheese_van 3y agoThank you for mentioning the "bat and ball" problem. Having neither a math nor CS background, I hadn't heard of it - and got it wrong. And reflecting on why I got it wrong I gained a little understanding of my own flawed mind. Why did I focus on a single variable and not a relationship? It set my mind wandering and was a nice morsel to digest with my breakfast. Thanks!
- richardw 3y agoOn backtracking, I thought tree-of-thought enabled that? "considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices" https://arxiv.org/abs/2305.10601 https://arxiv.org/abs/2305.10601 Generally with you though, this thing is not leading to real smarts and that's accepted by many. Yes, it'll fill in a few gaps with exponentially more compute but it's more likely that an algo change is required once we've maxed out LLM's.
- wbhart 3y agoYes, there are various approaches like tree-of-thought. They don't fundamentally solve the problem because there are just too many paths to explore and inference is just too slow and too expensive to explore 10,000 or 100,000 paths just for basic problems that no one wanted to solve anyway. The problem with solving such problems with LLMs is that if the solution to the problem is unlike problems seen in training, the LLM will almost every time take the wrong path and very likely won't even think of the right path at all. The AI really does need to understand why the paths it tried failed in order to get insight into what might work. That's how humans work (well, one of many techniques we use). And despite what people think, LLMs really don't understand what they are doing. That's relatively easy to demonstrate if you get an LLM off distribution. They will double down on obviously erroneous illogic, rather than learn from the entirely new situation.
- richardw 3y agoThank you for the thoughtful response
- insomagent 3y agoLet's say a model runs through a few iterations and finds a small, meaningful piece of information via "self-play" (iterating with itself without further prompting from a human.) If the model then distills that information down to a new feature, and re-examines the original prompt with the new feature embedded in an extra input tensor, then repeats this process ad-infinitum, will the language model's "prime directive" and reasoning ability be sufficient to arrive at new, verifiable and provable conjectures, outside the realm of the dataset it was trained on? If GPT-4,5,...,n can progress in this direction, then we should all see the writing on the wall. Also, the day will come where we don't need to manually prepare an updated dataset and "kick off a new training". Self-supervised LLMs are going to be so shocking.
- wbhart 3y agoPeople have done experiments trying to get GPT-4 to come up with viable conjectures. So far it does such a woefully bad job that it isn't worth even trying. Unfortunately there are rather a lot of issues which are difficult to describe concisely, so here is probably not the best place. Primary amongst them is the fact that an LLM would be a horribly inefficient way to do this. There are much, much better ways, which have been tried, with limited success.
- gmerc 3y agoAfter a year the entire argument you make boils down to “so far”.
- ra 3y agoIndeed. LLM is an application on a transformer trained with backpropagation. What stops you from adding a logic/mathematic "application" on the same transformer?
- seanhunter 3y agoNothing, and there are methods which allow these types of models to learn to use special purpose tools of this kind[1]. [1] https://arxiv.org/abs/2302.04761 https://arxiv.org/abs/2302.04761 Toolformer: Language Models Can Teach Themselves to Use Tools
- nullc 3y agoIt's also hard to know what the LLM has reasoned out vs has memorized. I like the very last example in my tongue-in-cheek article, https://nt4tn.net/articles/aixy.html https://nt4tn.net/articles/aixy.html Certainly the LLM didn't derive Fermat's theorem on sums of two squares under the hood (and, of course, very obviously didn't prove it correct-- as the code is technically incorrect for 2), but I'm somewhat doubtful that there was any function exactly like the template in codex's training set either (at least I couldn't quickly find any published code that did that). The line between creating something and applying a memorized fact in a different context is not always super clear.
- elliotec 3y agoHubris
- lucubratory 3y agoWhose, in this instance? I can see an argument for both
- est 3y ago> The reason LLMs fail at solving mathematical problems is because That's exactly what Go/Baduk/Weiqi players think some years ago. And superalignment is defintely OpenAI's major research objective: > https://openai.com/blog/our-approach-to-alignment-research https://openai.com/blog/our-approach-to-alignment-research > our AI systems are proposing very creative solutions (like AlphaGo’s move 37) When will mathematicians face the move 37 moment?
- Davidzheng 3y agoProbably in <3 years
- topspin 3y agoI don't know whether this particular article is bunk. I do know I've read many, many similar comments about how some complex task is beyond an conceivable model or system and then, years later, marveled at exactly that complex task being solved.
- jhanschoo 3y agoThe article isn't describing something that will happen years later, but now. The comment author is saying that this current model is not AGI as it likely can't solve university-level mathematics, and they are presumably open to the possibility of a model years down the line that can do that.
- nostrademons 3y agoWhat I wonder, as a computer scientist: If you want to solve grade school math problems, why not use an 'add' instruction? It's been around since the 50s, runs a billion times faster than an LLM, every assembly-language programmer knows how to use it, every high-level language has a one-token equivalent, and doesn't hallucinate answers (other than integer overflow). We also know how to solve complex reasoning chains that require backtracking. Prolog has been around since 1972. It's not used that much because that's not the programming problem that most people are solving. Why not use a tool for what it's good for and pick different tools for other problems they are better for? LLMs are good for summarization, autocompletion, and as an input to many other language problems like spelling and bigrams. They're not good at math. Computers are really good at math. There's a theorem that an LLM can compute any computable function. That's true, but so can lambda calculus. We don't program in raw lambda calculus because it's terribly inefficient. Same with LLMs for arithmetic problems.
- xwolfi 3y agoYou're missing the point: who's using the 'add' instruction ? You. We want 'something' to think about using the 'add' instruction to solve a problem. We want to remove the human from the solution design. It would help us tremendously tbh, just like I don't know, Google map helped me never to have to look for direction ever again ?
- marshray 3y agoWhen the solution requires arithmetic, one trick is to simply ask GPT to write a Python program to compute the answer. There's your 'add'.
- davidwritesbugs 3y agoInteresting, how do you use this idea? If you prompt the LLM "create a python Add function Foo to add a number to another number", "using Foo add 1 and 2", or somesuch, but what's to stop it hallucinating and saying "Sure, let me do that for you, foo 1 and 2 is 347. Please let me know if you need anything else."
- stephenboyd 3y agoDid they say it was an LLM? I didn’t see that in the reporting.
- 3cats-in-a-coat 3y agoI don't understand your thesis here it seems self-contradictory: 1. "I don't think this is real news / important because solving grade school math is not a predictor of ability to do complex reasoning." 2. "LLMs can't solve grade school math because they're bad at arithmetic, algebra and most importantly reasoning." So... from 2 automatically follows that LLMs with sufficiently better math may be sufficiently better at reasoning as you said "most importantly" reasoning is relevant for their ability to do math. Saying "most importantly reasoning" and then saying that reasoning is irrelevant if they can do math, is odd.
- poulpy123 3y agoI don't know for Q* of course, but all the tests I made with GPT4, and all what I've read and seen about it, show that it is unable to reason. It was trained with an unfathomable amount of data, so it can simulate reasoning very well, but it is unable to reason
- oezi 3y agoWhat is the difference between simulating reasoning very well and "actual" reasoning?
- parentheses 3y agoI think the poster meant that it's capable of having a high probability of correct reasoning - simulating reasoning is lossy, actual reasoning is not. Though, human reasoning is still lossy.
- silvaring 3y agoActual reasoning is made up of various biological feedback loops that happen in the body and brain, essentially your physical senses give you the ability to reason in the first place, without the eyes, ears etc there is no ability to learn basic reasoning, which is why kids who are blind or mute from birth have huge issues learning about object permanence, spatial awaraness etc. You cant expect human reasoning without human perception. My question is how does the AI perceive. Basically how good is the simulation for its perception. If we know that, then we can probably assess its ability to reason because we can compare it to the closest benchmark we have (your average human being). How do AI's see, how did they learn concepts in strings of words and pixels? How does the concept it learnt in text carry through to images of colors, of shapes? Does it show a transfer of conceptual understanding across both two and three dimentional shapes? I know these are more questions than answers, but its just things that I've been wondering about.
- ajuc 3y agoThis ship can't swim because only living creatures swim. It's true but it only shows your definition sucks.
- codedokode 3y ago> The reason LLMs fail at solving mathematical problems is because ...because they are too small and have too little weights. Cats cannot solve mathematical problems too, but unlike cats, neural network evolve.
- serf 3y ago>Cats cannot solve mathematical problems too, but unlike cats, neural network evolve. cats evolve plenty, pressure towards mathematical reasoning has stymied as of late what with the cans of food and humans.
- deleted 3y ago[deleted]
- CodeCompost 3y agoIt's a text generator that spits out tokens. It has absolutely no understanding of what it's saying. We as humans are attaching meaning to the generated text. It's the humans that are hallucinating, not the text generator.
- bottlepalm 3y agoThey've already researched this and have found model inside the LLM such as a map of the world - https://x.com/wesg52/status/1709551516577902782 https://x.com/wesg52/status/1709551516577902782. Understanding is key to how so much data can be compressed into a LLM. There really isn't a better way to store all of it better than plain understanding it.
- deleted 3y ago[deleted]
- jiggawatts 3y agoEverything you said about LLMs being "terrible at X" is true of the current generation of LLM architectures. From the sound of it, this Q* model has a fundamentally different architecture, which will almost certainly make some of those issues not terrible any more. Most likely, the Q* design is the very similar to the one suggested recently by one of the Google AI teams: doing a tree search instead of greedy next token selection. Essentially, current-gen LLMs predict a sequence of tokens: A->B->C->D, etc... where the next "E" token depends on {A,B,C,D} and then is "locked in". While we don't know exactly how GPT4 works, reading between the lines of the leaked info it seems that it evaluates 8 or 16 of these sequences in parallel, then picks the best overall sequence. On modern GPUs, small workloads waste the available computer power because of scheduling overheads, so "doing redundant work" is basically free up to a point. This gives GPT4 a "best 1 of 16" output quality improvement. That's great, but each option is still a linear greedy search individually. Especially for longer outputs the chance of a "mis-step" at some point goes up a lot, and then the AI has no chance to correct itself. All 16 of the alternatives could have a mistake in them, and now its got to choose between 16 mistakes. It's as if you were trying to write a maths proof, asked 16 students, and instructed them to not cooperate and write their proof left-to-right, top-to-bottom without pausing, editing, or backtracking in any way! It'd like to see how "smart" humans would be at maths under those circumstances. This Q* model likely does what Google suggested: Do a tree search instead of a strictly linear search. At each step, the next token is presented as a list of "likely candidates" with probabilities assigned to each one. Simply pick to "top n" instead of the "top 1", branch for a bit like that, and then prune based on the best overall confidence instead of the best next token confidence. This would allow a low-confidence next token to be selected, as long as it leads to a very good overall result. Pruning bad branches is also effectively the same as back-tracking. It allows the model to explore but then abandon dead ends instead of being "forced" to stick with bad chains of thought. What's especially scary -- the type of scary that would result in a board of directors firing an overly commercially-minded CEO -- is that naive tree searches aren't the only option! Google showed that you can train a neural network to get better at tree search itself, making it exponentially more efficient at selecting likely branches and pruning dead ends very early. If you throw enough computer power at this, you can make an AI that can beat the world's best chess champion, the world's best Go player, etc... Now apply this "AI-driven tree search" to an AI LLM model and... oh-boy, now you're cooking with gas! But wait, there's more: GPT 3.5 and 4.0 were trained with either no synthetically generated data, or very little as a percentage of their total input corpus. You know what is really easy to generate synthetic training data for? Maths problems, that's what. Even up to the point of "solve this hideous integral that would take a human weeks with pen and paper" can be bulk generated and fed into it using computer algebra software like Wolfram Mathematica or whatever. If they cranked out a few terabytes of randomly generated maths problems and trained a tree-searching LLM that has more weights than GPT4, I can picture it being able to solve pretty much any maths problem you can throw at it. Literally anything Mathematica could do, except with English prompting! Don't be so confident in the superiority of the human mind. We all thought Chess was impossible for computers until it wasn't. Then we all moved the goal posts to Go. Then English text. And now... mathematics. Good luck with holding on to that crown.
- kolinko 3y agoLLMs by themselves don’t learn from past past mistakes, but you could cycle inference steps and fine tuning/retraining steps. Also, you can store failed attempts and lessons learned in context.
- gmt2027 3y agoWe have an algorithm and computational hardware that will tune a universal function approximator to fit any dataset with emergent intelligence as it discovers abstractions, patterns, features and hierarchies. So far, we have not yet found hard limits that cannot be overcome by scaling the number of model parameters, increasing the size and quality of training data or, very infrequently, adopting a new architecture. The number of model parameters required to achieve a defined level of intelligence is a function of the architecture and training data. The important question is, what is N, the number of model parameters at which we cross an intelligence threshold and it becomes theoretically possible to solve mathematics problems at a research level for an optimal architecture that we may not yet have discovered. Our understanding does not extend to the level where we can predict N but I doubt that anyone still believes that it is infinity after seeing what GPT4 can do. This claim here is essentially a discovery that N may be much closer to where we are with today's largest models. Researchers at the absolute frontier are more likely to be able to gauge how close they are to a breakthrough of that magnitude from how quickly they are blowing past less impressive milestones like grade school math. My intuition is that we are in a suboptimal part of the search space and it is theoretically possible to achieve GPT4 level intelligence with a model that is orders of magnitude smaller. This could happen when we figure out how to separate the reasoning from the factual knowledge encoded in the model.
- waveBidder 3y agointelligence isn't a function unless you're talking about over every possible state of the universe.
- gmt2027 3y agoThere are well described links between intelligence and information theory. Intelligence is connected to prediction and compression as measures of understanding. Intelligence has nothing specific to do with The Universe as we known it. Any universe will do, a simulation, images or a set of possible tokens. The universe is every possible input. The training set is a sampling drawn from the universe. LLMs compress this sampling and learn the processes and patterns behind it so well that they can predict what should come next without any direct experience of our world. All machine learning models and neural networks are pure functions. Arguing that no function can have intelligence as a property is equivalent to claiming that artificial intelligence is impossible.
- lukego 3y ago> 1) they are terrible at arithmetic, 2) they are terrible at algebra The interaction can be amusing. Proving algebra non-theorems by cranking through examples until an arithmetic mistake finally leads to a "counter-example." It's like https://xkcd.com/882/ https://xkcd.com/882/ for theorems.
- greendesk 3y agoThinking is about associations and object visualisation. Surely a non-human system can build those, right? Pointing out only to a single product exposed to the public does not prove limitations for a theoretical limit.
- paulsutter 3y agoGood point. What would these AI people know about AI? You’re right, what they’re doing will never work You should make your own, shouldn’t take more than a weekend, right?
- d--b 3y agoYou make the asumption that Q* is a LLM, but I think OpenAI guys know very well that the current LLM architecture cannot achieve AGI. As the name suggests, this things is likely using some form of Q learning algorithm, which makes it closer to the DeepMind models than a transformer. My guess is that they pipe their LLM into some Q learnt net. The LLM may transform a natural language task into some internal representation that can then be handled by the Q-learnt model, which spits out something that can be transformed back again into natural language.
- jansan 3y agoThere is a paper about something called Q*. I have no idea if they are connected or if the name matched coincidentially. https://arxiv.org/abs/2102.04518 https://arxiv.org/abs/2102.04518
- wegfawefgawefg 3y agoThe real world is a space of continuous actions. To this day Q algorithms have been ones of discrete action outputs. I'd be surprised if a Q algorithm could handle the huge action space of language. Honestly its weird they'd consider the Q family. I figured we were done with that after PPO performed so well.
- wegfawefgawefg 3y agoAs an ML programmer, i think that approach sounds really too complicated. It is always a bad idea to render the output of one neural network into output space before feeding it into another, rather than have them communicate in feature space.
- vidarh 3y agoI feel very comfortable to say that while the ability to solve grade school maths is not a predictor of abilities at a research level, the advances needed to solve 1 and 2 will mean improving results across the board unless you take shortcuts (e.g. adding an "add" instruction as proposed elsewhere), because if you actually dig into prompting an LLM to follow steps for arithmetic what you quickly see is that problem has not been the ability to reason on the whole (that is not to suggest that the ability to reason is good enough), but ability to consistently and precisely follow steps a sufficient number of times. It's acting like a bored child who hasn't had following the steps and verifying the results repetitively drilled into it in primary school. That is not to say that their ability to reason is sufficient to reason at an advanced level yet, but so far what has hampered a lot of it has been far more basic. Ironically, GPT4 is prone to take shortcuts and make use of the tooling enabled for it to paper over its abilities, but at the same time having pushed it until I got it to actually do arithmetic of large numbers step by step, it seems to do significantly better than it used to at systematically and repetitively following the methods it knows, and at applying "manual" sanity checks to its results afterward. As for lemma conjecturing, there is research ongoing, and while it's by no means solved, it's also not nearly as dire as you suggest. See e.g.[1] That's not to suggest it's reasoning abilities are sufficient, but I also don't think we've seen anything to suggest we're anywhere close to hitting the ceiling of what current models can be taught to do, even before considering advancements in tooling around them, such as giving them "methods" to work to and a loop with injected feedback, access to tools and working memory. [1] https://research.chalmers.se/en/publication/537034 https://research.chalmers.se/en/publication/537034
- Tangokat 3y agoHow about this: - The Q* model is very small and trained with little compute. - The OpenAI team thinks the model will scale in capability in the same way the GPT models do. - Throwing (much) more compute at the model will likely allow it to solve research level math and beyond, perhaps also do actual logic reasoning in other areas. - Sam goes to investors to raise more money (Saudi++) to fund the extra compute needed. He wants to create a company making AI chips to get more compute etc. - The board and a few other OpenAI employees (notably Ilya) wants to be cautious and adopt a more "wait and see" approach. All of this is speculation of course.
- bambax 3y ago> 4) they (current LLMs) cannot backtrack when they find that what they already wrote turned out not to lead to a solution, and it is too expensive to give them the thousands of restarts they'd require to randomly guess their way through the problem if you did give them that facility This sounds like a reward function? If correctly implemented couldn't it enable an LLM to self-learn?
- ddalex 3y agoSpecifically what deep-Q learning (as in Q*?) does....
- two_in_one 3y agoThe reason LLMs solve school problems is because they've been trained on solutions. The problems are actually very repetitive. Not surprising for each 'new' of them there was something similar in training set. For research level problems there is nothing in training set. That's why they don't perform well. Just today I asked GPT4 a simple task. Having mouse position in zoomed and scrolled image find it's position in the original image. GPT4 happily wrote the code, but it was completely wrong. I had to fix it manually. However, the performance can be increased if there are several threads working on solution. Some suggesting and others analyzing the solution(s). This will increase the size of 'active' memory, at least. And decrease the load on threads, making them more specialized and deeper. This requires more resources, of course. And good management with task split. May be a dedicated thread for that.
- Enogoloyo 3y agoJust wait a little bit. You are not better than a huge GPU cluster with Monte Carlo search and computer verification for much longer. It will be more your job to find the interesting finds than doing the work of finding things in the first olace
- aremat 3y ago"A Mathematician" (Lenat and co.) did indeed attempt to approach creative theorem development from a radically different approach (syllogistic search-space exploration, not dissimilar to forward-chaining in Prolog), although they ran into problems distinguishing "interesting" results from merely true results: https://web.archive.org/web/20060528011654/http://www.comp.glam.ac.uk/pages/staff/efurse/Abstracts/Why-did-AM-halt.html https://web.archive.org/web/20060528011654/http://www.comp.g...
- Ldorigo 3y agoDid anyone claim that it would be a predictor of solving math problems at a research level? Inasmuch as we can extrapolate from the few words in the article it seems more likely that the researchers working on this project identified some emergent reasoning abilities exemplified with grade level math. Math literacy/ability that is comparable to the 0.1% of humans is not the end goal of OpenAI, "general intelligence" is. I have plenty of people in my social circle who are most certainly "generally intelligent" yet have no hope of attaining those levels of mathematical understanding. Also note that we don't know if Q* is just a "current LLM" (with some changes)
- afpx 3y agoWhat amazes me is how close it gets to the right answer, though. Pick a random 10-digit number, then ask the next 20 numbers in sequence. I feel like the magic in these LLMs is in how they work well in stacks, trees or in seqence. They become elements of other data structures. Consider a network of these, combined with other specialized systems and an ability to take and give orders. With reinforcement learning, it could begin building better versions of itself.
- naikrovek 3y agoit's always nice to see HN commenters with so much confidence in themselves that they feel they know a situation better than the people who are actually in the situation being discussed. Do you really believe that they don't have skilled people on staff? Do you really believe that your knowledge of what OpenAI is doing is a superset of the knowledge of the people who work at OpenAI? give me 0.1% of your confidence and I would be able to change the world.
- mycologos 3y agoThe people inside a cult are not the most trustworthy sources for what the cult is doing.
- nopromisessir 3y agoThis is defamatory and unfounded. OpenAI is exploring the limits of computation. The more I see this kind of unfounded slander, the more confident I become that this outfit might be the most important in the face of planet earth. So many commentors here are starting to sound like priests of the Spanish inquisition. Do you seriously expect a community of technologist and science advocates to be fearful of such assertions without evidence? It's a waste of breath. All credibility just leaves the room instantly.
- naikrovek 3y agoit is very odd that you consider OpenAI a cult. go spend time in a cult, then come back and tell us how much of a cult OpenAI is. you know nothing
- BenoitP 3y agoWhat do you think of integrating propositional logic, first order logic and sat solvers in LLM output? ie forcing each symbol an LLM outputs to have its place in a formal proposition. And getting a prompt from the user to force that some parts be satisfiable. I know this is not how us humans craft our thoughts, but maybe an AI can optimize to death the conjunction of these tools. The LLM just being an universal API to the core of formal logic.
- EGreg 3y agoAs someone who studied math in grad school as part of a PhD program, worked at a hedge fund and went on to work on software and applied math, I call bullshit on this. Math and Logic is just low-dimensional symbol manipulation that computers can easily do. You can throw data at them and they’ll show you theories that involve vectors of 42,000 variables while Isaac Newton had 4 and Einstein had 7 with Levi-Civita calculus. In short, what you consider “reasoning”, while beautiful in its simplicity, is nevertheless crude approximations to complex systems, such as linear regression or least squares. 3 days ago AI predicted fluid dynamics better than humans: https://www.sciencedaily.com/releases/2023/11/231120170956.htm https://www.sciencedaily.com/releases/2023/11/231120170956.h... Google’s AI predicts weather now faster and better than Current systems built by humans: https://www.zdnet.com/google-amp/article/ai-is-outperforming-our-best-weather-forecasting-tech-thanks-to-deepmind/ https://www.zdnet.com/google-amp/article/ai-is-outperforming... AlphaZero based on MCTS years ago beat Rybka and all human-built systems in chess: https://www.quora.com/Did-AlphaZero-really-beat-Stockfish https://www.quora.com/Did-AlphaZero-really-beat-Stockfish And it can automate science and send it into overdrive: https://www.pbs.org/newshour/amp/science/analysis-how-ai-is-helping-astronomers-study-the-universe https://www.pbs.org/newshour/amp/science/analysis-how-ai-is-...
- sheepscreek 3y ago> They learn from failed attempts in ways that LLMs do not seem to. The material they are trained on surely contributes to this problem. For transformer models, they do learn from their mistakes but only during the training stage. There’s no feedback loop during inference, and perhaps there needs to be something; like real-time fine-tuning.
- yallneedtoget 3y ago[dead]
- wegfawefgawefg 3y agoYour comment is regarding LLMs, but Q* may not refer to an LLM. As such, our intuition about the failure of LLM's may not apply. The name Q* likely refers to a deep reinforcement learning based model. To comment, in my personal experience, reinforcement learning agents learn in a more relatable human way than traditional ml, which act like stupid aliens. RL Agents try something a bunch of times, mess up, and tweak their strategy. After some extreme level of experience, they can make wider strategic decisions that are a little less myopic. RL agents can take in their own output, as their actions modify the environment. RL Agents also modify the environment during training, (which I think you will agree with me is important if you're trying to learn the influence of your own actions as a basic concept). LLM's, and traditional ml in general, are never trained in a loop on their own output. But in DRL, this is normal. So if RL is so great and superior to traditional ml why is RL not used for everything? Well the full time horizon that can be taken into consideration in a DRL Agent is very limited, often a handful of frames, or distilled frame predictions. That prevents them from learning things like math. Traditionally RL bots have been only used for things like robotic locomotion, chess, go. Short term decision making that is made given one or some frames of data. I don't even think any RL bots have learned how to read english yet lol. For me, as a human, my frame predictions exist on the scale of days, months, and years. To learn math I've had to sit and do nothing for many hours, and days at a time, consuming my own output. For a classical RL bot, math is out of the question. But, my physical actions, for ambulation, manipulation, and balance, are made for me by specialized high speed neural circuits that operate on short time horizons, taking in my high level intentions, and all the muscle positions, activation, sensor data, etc. Physical movement is obfuscated from me almost in entirety. (RL has so far been good at tasks like this.) With a longer frame horizon, that predicts frames far into the future, RL can be able to make long term decisions. It would likely take a lifetime to train. So you see now why math has not been accomplished by RL yet, but I don't think the faculty would be impossible to build into an ml architecture. An RL bot that does math would likely spin on its own output for many many frames, until deciding that it is done, much like a person.
- jug 3y ago1. OpenAI researchers used loaded and emotional words, implying shock or surprise. It's not easy to impress an OpenAI researcher like this, and above all, they understand the difficulty difference between teaching AI grade school and complex math since many years. They also understand that solving math with any form of reliability is only an emergent property in quite advanced LLM's. 2. Often, research is made on toy models and if this would be such a model, acing grade school problems (as per the article) would be quite impressive to say the least as this ability simply isn't emergent early in current LLM's. What I think might have happened here is a step forward in AI capacity that has surprised researchers not because it is able to do things it couldn't at all do before, but how _early_ it is able to do so.
- theonlybutlet 3y agoYou're underestimating the power of LLM's. I'll address two of your points as the other two stem from this. They can't backtrack that's purely just design and can be easily trained there's no need to simulate at random until it gets the answer, if allowed to review it's prior answers and consider this, if often can reason a better answer. Further more breaking down problems. This is easily demonstrated when looking at how accuracy improves when you ask it to explain it's reasoning as it calculates (break it down into smaller problems). The same for humans, large mathematical problems are solved using learned methods to breakdown and simplify calculations into those easier for us to calculate and build up. If the model was able to self adjust weightings based on it's finding this would further improve it (another design limitation we'll eventually get to improve, reinforcement learning). Much like 2+2=4 is your instantaneous answer, the neural connection has been made so strong in our brains by constant emphasis we no longer need to think of an abacus each time we get to the answer 4. You're also ignoring the emergent properties of these LLMs, theyre obviously not yet at human level but they do understand the underlying values and can reason using this value. Semantic search/embeddings is evidence of this.
- arendtio 3y agoTo some degree you are right, but I think you forget, that the things they solved already (talking and reasoning about a world that was only presented in the form of abstractions (words)) were supposed to be harder than having a good understanding of numbers, quantities, and logical terms. My guess is, that they saw the problems ChatGPT has today and worked on solving those problems. And given how important numbers are, they tried to fix how ChatGPT handles/understands numbers. After doing that, they saw how this new version performed much better and predicted, that further work in this area could lead to solving real-world math problems. I don't think that we will be presented with the highway to singularity, but probably one more step in that direction.