12 ms·
>Today's AI models are missing the ability to reason abstractly, including asking and answering questions of "Why?" and "How?" This claim seems over general, b
by fergal_reid 3y ago
>Today's AI models are missing the ability to reason abstractly, including asking and answering questions of "Why?" and "How?"
This claim seems over general, because you can ask gpt-4 'Why' and 'How' questions and it seems to do a pretty good job.
The author doesn't provide a lot of contrary evidence.
There's so many articles saying "LLMs can't do X" that leave me wondering whether the author has even tried. Maybe they've tried and have some more sophisticated argument, but I often don't see it.
If I was going to knock LLMs for being unable to do basic science, in particular, I'd make sure to do some experiments first!
- famouswaffles 3y agoThe problem is that today's state of the art is far too good for low hanging fruit. There isn't a testable definition of GI that GPT-4 fails that a significant chunk of humans wouldn't also fail so you're often left with weird ad-hominins ("Forget what it can do and results you see. It's "just" predicting the next token so it means nothing") or imaginary distinctions built on vague and ill defined assertions ( "It sure looks like reasoning but i swear it isn't real reasoning. What does "real reasoning" even mean ? Well idk but just trust me bro")
- timmytokyo 3y ago>There isn't a testable definition of GI [...] This to me is the fundamental issue in discussions and debates about LLMs. Despite assertions by some psychologists (who themselves are practitioners of perhaps the fuzziest of "sciences"), intelligence is an entirely nebulous concept. Everyone means something different when they use the word. I can think of no better illustration of the problem than the authors of the "Sparks of AGI" paper resorting to a definition of intelligence presented in the Wall Street Journal of all places. That the WSJ definition was part of an editorial defending the Bell Curve is just the cherry on top.
- janalsncm 3y agoDo you know what their definition was by any chance? And yes, a cursory glance at the Wikipedia page for intelligence shows there’s no one agreed upon definition of intelligence. A more useful framing is to say we’re not creating “intelligence” per se but automating tasks. GPT4 is an automated writer. Stable diffusion is an automated image creator. Alpha Go was an automated Go player. Google search automates the work of a reference librarian. With that in mind, it’s immediately obvious how much of a waste of time it is to argue whether ChatGPT is “intelligent” or not. Who cares. What we are doing is automating all of the things which brains used to do.
- TheOtherHobbes 3y agoOne problem is that academic CS-researcher intelligence is completely different to average human intelligence. Maybe 5% of the population can learn how to solve partial differential equations. Virtually all of the population can manage extended family-related conversations over Christmas. Even when drunk. Human intelligence is mostly social, and mostly not scientific. The average human is incredibly bad at model building and self-correcting prediction. What actually happens is that humans have developed a kind of collective cultural exoskeleton which protects - more or less - from the consequences of poor choices. But it doesn't take much for that to stop working. Covid denial and climate change denial are just two examples. The cost if living in this space is having to learn a lot of heavily scripted cues. There's a long list of acceptable and unacceptable behaviours and social registers in different social situations. It varies by culture. But generally humans can navigate this space without thinking too hard about it. Academic intelligence is completely different. There's long been a joke that an AI researcher's ideal intelligent system is another AI researcher, with typical AI researcher interests - math, puzzles, abstract language models, music in an engineering way, and so on. Current LLMs are the first cross-over product which shows signs of moving into the first space from the second. You can imagine a future system which uses facial and gait profiling to read emotions, and links a tokenised language model with a tokenised model of various transitions through emotional and social states. Personal background will be missing, and that's not hard to invent. And now you have something that mimics a large part of social intelligence. Only it has the potential to do it better than humans do.
- janalsncm 3y ago> It's "just" predicting the next token so it means nothing This form of argument should raise red flags for everyone. It is an argument against the possibility of emergence, that a sufficient number of simple systems cannot give rise to more complex ones. Human beings are “just” a collection of cells. Calculators are “just” a stupid electric circuit. The fact is, putting basic components together is the only way we know how to make things. We can use those smaller component to make a more complex thing to accomplish a more complex task. And emergence is everywhere in nature as well.
- mysterydip 3y agoI don't have access to any so I can't ask to check, but are they able to complete basic math problems (numeric or word)? I know earlier ones would say like 5*7 was 96 etc
- bcherny 3y agoGPT4: > What is 5*7? > 5 multiplied by 7 equals 35. > Why? > Multiplication is a process of adding a number to itself a certain number of times. When you multiply 5 by 7, it's the same as saying "add 5 to itself 7 times." If you add up seven 5s (5 + 5 + 5 + 5 + 5 + 5 + 5), you get 35. That's why 5 multiplied by 7 equals 35. > (7 “why?”s later…) > The ultimate "why" behind many fundamental rules and concepts, including those in mathematics, can often circle back to human needs and the desire to understand and interact effectively with our world. > As humans, we've developed systems and structures, such as mathematics, to organize, interpret, and predict our surroundings. These systems help us in practical tasks like trading goods, constructing buildings, and planning time, but also in abstract reasoning and problem solving. The rules within these systems, such as the operation of multiplication in mathematics, exist because they are useful, they work consistently, and they can be universally understood and applied. > So, why does 5 multiplied by 7 equal 35, and why does multiplication work the way it does? Because humans, over many generations, found that this system was a powerful and reliable tool for interacting with the world and with each other.
- jacquesm 3y agoThat could still be explained by that example having been part of the training set so it knows how to give the right answer. But there must be more going on than that: The square root of 232444232 is approximately 15229. (which is wrong, it is approximately 15246.12) The sum of 2341347345 and 234823542354 is 237165889699. Which is the right answer. So there may be some special casing happening there.
- bcherny 3y agoI mean, I don’t know the square root of 232444232 off the top of my head either..
- kec 3y agoImprecise language, but the article is specifically referring to questions like “why do you think I asked you that?” Or “how are you answering these questions?”. LLM’s can’t engage with these types of questions, the best they can do is to regurgitate a canned response peppered with some prompt history.
- kenjackson 3y agoIn fairness most humans can’t either. Try going to a random person at the park and asking them “Explain the relationship between Romeo and Juliet and Star Trek”. And then ask them why they think you asked that question. They’ll mostly be befuddled I suspect.
- almost_usual 3y agoSo knowledge or memorization of culture is intelligence? What if that personal steals your wallet without you being aware while you ask them that question because they need food. Is that intelligent?
- robbomacrae 3y agoI had to go and try this exact line of questioning with ChatGPT because I suspected this might lead to a weakness in it not admitting when it just doesn't have a clue (which would have been my answer)... mind you its a big human weakness/tendency to not admit lack of knowledge. But the answer was surprisingly candid and yet thoughtful: """ I can't know for sure why you asked the question about the relationship between "Romeo and Juliet" and "Star Trek," as I don't have access to your personal thoughts or context. However, some potential reasons might include: Academic Inquiry: You might be exploring themes in literature or media studies and are interested in drawing connections between different works across genres and time periods. Creative Inspiration: If you're a writer, artist, or content creator... """ There were some others but overall I thought the initial disclaimer along with some possible theories approach was spot on and a lot better than my "no clue" knee jerk reaction.
- ryanjshaw 3y agoAlso from the article: > What makes human intelligence different from today's AI is the ability to ask why, reason from first principles, and create experiments and models for testing hypotheses. This is quite unfair. The AI doesn't have I/O other than what we force-feed it through an API. Who knows what will happen if we plug it into a body with senses, limbs, and reproductive capabilities? No doubt somebody is already building an MMORPG with human and AI characters to explore exactly this while we wait for cyborg part manufacturing to catch up.
- version_five 3y agoThis is just wrong, it has no external goals, it just predicts next tokens or behaves in some other way that has minimized a training loss. It doesn't matter what you "plug it in to", it will just do what you tell it. You could speculate there might be instructions that lead to emergent behavior, but then your back to just speculating about how AI might work. Current llms don't work the way you're implying.
- Philadelphia 3y agoIt also can’t learn. Once the training is done, the network is set in stone.
- adamisom 3y agoMakes me wonder why we don’t see deployed models that keep learning during inference.
- LegitShady 3y agoMicrosoft tay has entered the chat
- Der_Einzige 3y agoThe curse of dimensionality and exploding/vanishing gradients are why incremental learning is still so rare.
- dontmobile 3y agoThere are limitations with LLMs but nobody is being clear about it. The overall state of LLMs can be distilled into 3 points: 1. LLMs Can produce output that is equal in intelligence and creativity to humans. It can even produce output that is objectively better than humans. This EVEN applies to novel responses that are completely absent from the training set. This is the main reason why there's so much hype around LLMs right now. 2. The main problem is that LLMs can't produce good output consistently. Sometimes the output is better, sometimes it's the same, sometimes it's the worse. LLMs sometimes "hallucinate", they are sometimes inconsistent, they have an obvious memory problems. But none of these problems completely preclude the LLM from being able to produce output that is objectively better or the same as human level reasoning... it's just not doing this consistently. 3. Nobody fully understands the internal state of LLMs. We have limited understanding of what's going on here. We can understand inputs and outputs but the internal thought process is not completely understood. Thus we can only make limited statements about how an LLM thinks. Nobody can make a statement that LLMs obviously have zero understanding of the world, nobody can make a statement that LLMs are just stochastic parrots because we don't really get whats going on internally. We only have output from LLMs that are remarkably novel and intelligent and output from LLMs that are incredibly stupid and inconsistent. The data does not point towards a definitive conclusion, it only points towards possibilities. There's actually a cargo cult around downplaying AI. There are people who say clearly the AI is a stochastic parrot and they point to the intention of the algorithm itself behind the LLM. Yes the algorithm at the lowest level can be thought of as a next text predictor. But this is just a low level explanation. It's like saying a computer system is simply a turing machine executing simplistic instructions from a tape roll when such instructions can form things like games and 3D simulations of entire open worlds. The high level characteristics of this AI is something we currently cannot understand. Yes we built a text predictor, but something else that was not expected came out as an emergent property and this emergent property is something we still cannot make a definitive statement about. What does the future hold? What follows is my personal opinion on this matter: I believe we will never be able to make a definitive statement about LLMs or even AGI. We will never be able to fully understand these things and instead AGI will come about from a series of trials, errors and accidents. What we build will largely come about as an art and as unexpected emergent properties of trying different things. I believe this for two reasons. The first reason is philosophical. There's this sort of blurry concept that I believe that a complex intelligence cannot fully comprehend something that is equal in complexity to itself. We can only partially understand complexity equal to ourselves by symbolically abstracting parts away but not everything can be abstracted like this. Sometimes true understanding involves comprehension of the entire complex crystal without abstracting any part of it away. I believe that the concept of "intelligence" is such a crystal, but that's just a guess. The second reason is scientific. We've had physical creations of complex intelligence right in front of ours eyes that we can touch, manipulate and influence for decades. The human brain and other animal brains have been studied extensively and our understanding has been consistently far away from any form of true understanding. Given the evidence of the failure to understand the human brain even when it's right in front of us, I'd say we're unlikely to ever completely understand LLMs as well.
- jmh117 3y agoIsn't that you asking the Whys and How's? If you asked an LLM "What's 5*4?" and it responded with "Why do you want to know that?", the LLM would be doing the abstract reasoning.
- LegitShady 3y agoNo, those would simply be the most statistically likely words given it's training set and input. It has no idea what 5'4" is to do abstract reasoning. It's a statisitic word probability model not an abstract thought model. They are stochastic parrots with a large complex training set, not reasoning.
- deleted 3y ago[deleted]
- paganel 3y agoI did try the Google LLM thing, Bard I think it is called, about the result of a football match that has marked the sporting history of my country (Romania). According to Bard we did manage to defeat the Swedes by two goals to one back at the 1994 Euro Championships, which, to put it bluntly, is pretty damn far from the truth (the Swedes managed to go through to the World Cup semifinals after winning on penalty shoot-outs, the score had been 2-2 after 120 minutes). I didn’t make any further inquiries, suffice is to say that there’s no “intelligence” in the concept of LLMs to speak of as long as it can’t even correctly answer a question that non-smart tech had been able answer correctly for years.
- mmq 3y agoYou used the model for fact checking. These models are not good at being used as a knowledge base.
- jacquesm 3y agoI would never use an LLM for fact checking, then you'd have to check again using something else.
- mmq 3y agoUsually for asking questions about specific details, people are using RAG (Retrieval Augmented Generation) to ground the information and provide enough context for the llm to return the correct answers. This means additional engineering plumbing and very specific context to query information from.
- nomel 3y agoFact recollection is not most people’s definition of intelligence. In fact, it’s something that the only known intelligent systems are infamously bad at.
- paganel 3y agoSo you’re saying I used it wrong? How does that help the pro-LLM case? What should have I asked it? Some philosophical question that didn’t involve “fact recollection”? At least this latest tech bluff is not bankrupting regular people like the crypto tech bluff had done.
- meling 3y agoTo be a scientist, the LLM should be asking fundamental questions (define hypotheses) on its own without human input and try to come up with answers.
- TeMPOraL 3y agoHuman scientists don't spontaneously grow on trees, they're being taught to ask such questions. LLMs could be too.
- YeGoblynQueenne 3y agoThe article: >>Today's AI models are missing the ability to reason abstractly, including asking and answering questions of "Why?" and "How?" Your comment: >> This claim seems over general, because you can ask gpt-4 'Why' and 'How' questions and it seems to do a pretty good job. The article says today's AI models can't ask why and how. You say _you_ can ask why and how.