7 ms·
When you look at how humans play chess they employ several different cognitive strategies. Memorization, calculation, strategic thinking, heuristics, and learn
by CSMastermind 2y ago
When you look at how humans play chess they employ several different cognitive strategies. Memorization, calculation, strategic thinking, heuristics, and learned experience.
When the first chess engines came out they only employed one of these: calculation. It wasn't until relatively recently that we had computer programs that could perform all of them. But it turns out that if you scale that up with enough compute you can achieve superhuman results with calculation alone.
It's not clear to me that LLMs sufficiently scaled won't achieve superhuman performance on general cognitive tasks even if there are things humans do which they can't.
The other thing I'd point out is that all language is essentially synthetic training data. Humans invented language as a way to transfer their internal thought processes to other humans. It makes sense that the process of thinking and the process of translating those thoughts into and out of language would be distinct.
- PaulDavisThe1st 2y ago> It's not clear to me that LLMs sufficiently scaled won't achieve superhuman performance on general cognitive tasks If "general cognitive tasks" means "I give you a prompt in some form, and you give me an incredible response of some form " (forms may differ or be the same) then it is hard to disagree with you. But if by "general cognitive task" you mean "all the cognitive things that human do", then it is really hard to see why you would have any confidence that LLMs have any hope of achieving superhuman performance at these things.
- jhrmnn 2y agoEven in cognitive tasks expressed via language, something like a memory feels necessary. At which point it’s not a LLM as in a generic language model. It would become a language model conditioned on the memory state.
- ddingus 2y agoMore than a memory. Needs to be a closed loop, running on its own. We get its attention, and it responds, or frankly if we did manage any sort of sentience, even a simulation of it, then the fact is it may not respond. To me, that is the real test.
- nox101 2y agoIt sounds like you think this research is wrong? (it claims llms can not reason) https://arstechnica.com/ai/2024/10/llms-cant-perform-genuine-logical-reasoning-apple-researchers-suggest/ https://arstechnica.com/ai/2024/10/llms-cant-perform-genuine... or do you maybe think no logical reasoning is needed to do everything a human can do? Tho humans seem to be able to do logical reasoning
- astrange 2y agoIt says "current" LLMs can't "genuinely" reason. Also, one of the researchers then posted an internship for someone to work on LLM reasoning. I think the paper should've included controls, because we don't know how strong the result is. They certainly may have proven that humans can't reason either.
- mannykannot 2y agoIf they had human controls, they might well show that some humans can’t do any better, but based on how they generated test cases, it seems unlikely to me that doing so would prove that humans cannot reason (of course, if that’s actually the case, we cannot trust ourselves to devise, execute and interpret these tests in the first place!) Some people will use any limitation of LLMs to deny there is anything to see here, while others will call this ‘moving the goalposts’, but the most interesting questions, I believe, involve figuring out what the differences are, putting aside the question of whether LLMs are or are not AGIs.
- CSMastermind 2y agoThe later. While I generally do suspect that we need to invent some new technique in the realm of AI in order for software to do everything a human can do, I use analogies like chess engines to caution myself from certainty.
- bbor 2y agoI’ll pop in with a friendly “that research is definitely wrong”. If they want to prove that LLMs can’t reason, shouldn’t they stringently define that word somewhere in their paper? As it stands, they’re proving something small (some of today’s LLMs have XYZ weaknesses) and claiming something big (humans have an ineffable calculator-soul). LLMs absolutely 100% can reason, if we take the dictionary definition; it’s trivial to show their ability to answer non-memorized questions, and the only way to do that is some sort of reasoning. I personally don’t think they’re the most efficient tool for deliberative derivation of concepts, but I also think any sort of categorical prohibition is anti-scientific. What is the brain other than a neural network? Even if we accept the most fringe, anthropocentric theories like Penrose & Hammerhoff’s quantum tubules, that’s just a neural network with fancy weights. How could we possibly hope to forbid digital recreations of our brains from “truly” or “really” mimicking them?
- threeseed 2y ago> It's not clear to me that LLMs sufficiently scaled won't achieve superhuman performance To some extent this is true. To calculate A + B you could for example generate A, B for trillions of combinations and encode that within the network. And it would calculate this faster than any human could. But that's not intelligence. And Apple's research showed that LLMs are simply inferring relationships based on the tokens it has access to. Which you can throw off by adding useless information or trying to abstract A + B.
- Dylan16807 2y ago> To calculate A + B you could for example generate A, B for trillions of combinations and encode that within the network. And it would calculate this faster than any human could. I don't feel like this is a very meaningful argument because if you can do that generation then you must already have a superhuman machine for that task.
- shkkmo 2y agoSure, when humans use multiple skill to address a specific problem, you can sometimes outperform them by scaling a spefic one of those skills. When it comes to general intelligence, I think we are trying to run before we can walk. We can't even make a computer with a basic, animal level understanding of the world. Yet we are trying to take a tool that was developed on top of system that already had an understanding of the world and use it to work backwards to give computers an understanding of the world. I'm pretty skeptical that we're going to succeed at this. I think you have to be able to teach a computer to climb a tree or hunt (subhuman AGI) before you can create superhuman AGI.
- senand 2y agoThis seems quite reasonable, but I recently heard a podcast (https://www.preposterousuniverse.com/podcast/2024/06/24/280-francois-chollet-on-deep-learning-and-the-meaning-of-intelligence/ https://www.preposterousuniverse.com/podcast/2024/06/24/280-...) that LLMs are more likely to be very good at navigating what they have been trained on, but very poor at abstract reasoning and discovering new areas outside of their training. As a single human, you don't notice, as the training material is greater than everything we could ever learn. After all, that's what Artificial General Intelligence would at least in part be about: finding and proving new math theorems, creating new poetry, making new scientific discoveries, etc. There is even a new challenge that's been proposed: https://arcprize.org/blog/launch https://arcprize.org/blog/launch > It makes sense that the process of thinking and the process of translating those thoughts into and out of language would be distinct Yes, indeed. And LLMs seem to be very good at _simulating_ the translation of thought into language. They don't actually do it, at least not like humans do.
- klabb3 2y ago> As a single human, you don't notice, as the training material is greater than everything we could ever learn. This bias is real. Current gen ai works proportionally well the more known it is. The more training data, the better the performance. When we ask something very specific, we have the impression that it’s niche. But there is tons of training data also on many niche topics, which essentially enhances the magic trick – it looks like sophisticated reasoning. Whenever you truly go “off the beaten path”, you get responses that are (a) nonsensical (illogical) and (b) “pulls” you back towards a “mainstream center point” so to say. Anecdotally of course.. I’ve noticed this with software architecture discussions. I would have some pretty standard thing (like session-based auth) but I have some specific and unusual requirement (like hybrid device- and user identity) and it happily spits out good sounding but nonsensical ideas. Combining and interpolating entirely in the the linguistic domain is clearly powerful, but ultimately not enough.
- NemoNobody 2y agoWhat part of AI today leads you to believe that an AGI would be capable of self directed creativity? Today that is impossible - no AI is truly generating "new" stuff, no poetry is constructed creatively, no images are born from a feeling, inspiration is only part of AI generation is you consider it utilizing it's training data, which isn't actually creativity. I'm not sure why everyone assumes an AGI would just automatically do creativity considering most people are not very creative, despite them quite literally being capable, most people can't create anything. Why wouldn't an AGI have the same issues with being "awake" that we do? Being capable of knowing stuff - as you pointed out, far more facts than a person ever could, I think an awake AGI may even have more "issues" with the human condition than us. Also - say an AGI comes into existence that is awake, happy and capable of truly original creativity - why tf does it write us poetry? Why solve world hunger - it doesn't hunger. Why cure cancer - what can cancer do to to it? AGI as currently envisioned is a mythos of fantasy and science fiction.
- TheOtherHobbes 2y agoChess is essentially a puzzle. There's a single explicit, quantifiable goal, and a solution either achieves the goal or it doesn't. Solving puzzles is a specific cognitive task, not a general one. Language is a continuum, not a puzzle. The problem with LLMs is that testing has been reduced to performance on language puzzles, mostly with hard edges - like bar exams, or letter counting - and they're a small subset of general language use.