10 ms·
The author makes this assertion about LLMs rather casually: >They don’t engage in logical reasoning. This is still a hotly debated question, but at this point
by superbatfish 2y ago
The author makes this assertion about LLMs rather casually:
>They don’t engage in logical reasoning.
This is still a hotly debated question, but at this point the burden of proof is on the detractors. (To put it mildly, the famous "stochastic parrot" paper has not aged well.)
The claim above is certainly not something that should be stated as fact to a naive audience (i.e. the authors' intended audience in this case). Simply asserting it as they have done -- without acknowledging that many experts disagree -- undermines the authors' credibility to those who are less naive.
- cristiancavalli 2y agoDisagree — proponents of this point still have yet to prove reasoning and other studies agree about “reasoning” being potentially fake/simulated: https://the-decoder.com/apple-ai-researchers-question-openais-claims-about-o1s-reasoning-capabilities/ https://the-decoder.com/apple-ai-researchers-question-openai... Just claiming a capability does not make it true and we have 0 “proof” of original reasoning that can be proved coming from these models. Especially given the potential cheating in current SOTA benchmarks
- ninetyninenine 2y agoIt’s stupid. You can prove that LLMs can reason by simply giving it a novel problem where no data exists and having it solve that problem. LLMs CAN reason. Whether it can’t reason is not provable. To prove that you have to give the LLM every possible prompt that it has no data for and effectively show it never reasons and gets it wrong all the time. Not only is the proof impossible but it’s already been falsified as we have demonstrable examples of LLMs reasoning. Literally I invite people to post prompts and correct answers to ChatGPT where it is trivially impossible for that prompt to exist in the data. Every one of those examples falsifies the claim that LLMs can’t reason. Saying LLMs can’t reason is an overarching claim similar to the claim that humans and LLMs always reason. Humans and LLMs don’t always reason. But they can reason.
- Miraste 2y agoAnswering novel prompts isn't proof of reasoning, only pattern matching. A calculator can answer prompts it's never seen before too. If anything, I would come down on the reasoning side, at least for recent CoT models-but it's not a trivial question at all.
- cristiancavalli 2y agoThis is a fun thought experiment and made me reminisce on my Epistemology classes — something I think the current AI conversation would benefit greatly from. I’m super excited about what we’ve created here — less from the practical standpoint and more from a philosophical one where we get to interact with another form of distilled knowledge. It’s really too bad so much is breathless hype and grift because the philosophy student in me just wants to bask in thinking about this different form/medium/distillation of knowledge we now get to interact with. Comments like these help to reinvigorate that love though so thank you!
- radlad 2y agoAre there any good Epistemology resources online? Seems like we could all benefit from this these days.
- cristiancavalli 2y agoI actually just sat down to crack open MITs Theory of Knowledge and it seems promising and free: https://ocw.mit.edu/courses/24-211-theory-of-knowledge-spring-2014/ https://ocw.mit.edu/courses/24-211-theory-of-knowledge-sprin... This also looks promising: https://hiw.kuleuven.be/en/study/prospective/OOCP/introduction-to-epistemology https://hiw.kuleuven.be/en/study/prospective/OOCP/introducti... If you wanted something a bit different Wittgenstein’s Tractatus has always made my head spin with possibilities: https://people.umass.edu/klement/tlp/tlp-hyperlinked.html https://people.umass.edu/klement/tlp/tlp-hyperlinked.html
- ninetyninenine 2y ago
- hnthrow90348765 2y ago>Disagree — proponents of this point still have yet to prove reasoning and other studies agree about “reasoning” being potentially fake/simulated: https://the-decoder.com/apple-ai-researchers-question-openai https://the-decoder.com/apple-ai-researchers-question-openai... ??? https://the-decoder.com/language-models-use-a-probabilistic-version-of-genuine-reasoning/ https://the-decoder.com/language-models-use-a-probabilistic-...
- cristiancavalli 2y agoYes people are claiming different things yet no definitive proof has been offered given the varying findings. I can cite another 3 papers which agree with my point and you can probably cite just as many if not more supporting yours. I’m arguing against people depicting what is not a forgone conclusion as such. It seems like in people’s rush to confirm their own preconceived notions people forget that, although a theory may be convincing, it may not be true. Evidence in this very thread of a well-known SOTA LLM not being able to tell which is greater between two numbers indicates to me that what is being called “reasoning” is not what humans do. We can make as many excuses we want per the tokenizer or whatever but then forgive me for not buying the super or even general “intelligence” of this software. I still like these tools though, even if I have to constantly vet everything they say as they often tend to just outright lie, or perhaps more accurately: repeat lies in their training data even if you can elicit a factual response on the same topic.
- semiquaver 2y agoWhat would definitive proof look like? Can you definitively prove that your brain is capable of reasoning and not a convincing simulation of it?
- cristiancavalli 2y agoI can’t and that’s pretty cool to think about! Of course if we’re going that far down the chain of assumption we’re not quite ready to talk about LLMs imo (then again maybe it would be the perfect place to talk about them as contrast/comparison; certainly exciting ideas in that light). From my own perspective: if we’re gonna say these things reason and we’re using the definition of reasoning we apply to humans, then being able to reason through the trivial cases they fail to today would be a start. To the proponents of “they reason sometimes but not others” my question is why? What reason does it have to not reason and why if it is reasoning it still fails on trivial things that are variations of its own training data? I would also expect that these models would use reasoning to find new things like humans do but without humans essentially guiding the model to the correct awnser or the model just brute-forcing a problem-space with a set of rules/heuristics. Not exhaustive but a good start I think. These models have trouble currently even doing the advertised things like “book a trip for me” once a UI update happens so I think it’s a great indication we don’t quite have the intelligence/reasoning aspect worked out. Another question I have: would a form of authentic reasoning in a model give rise to a model having an aesthetic? Could this be some sort of indicator of having created a “model of the world”? Does the model of the world perhaps imply a value judgement about it given that if one was super intelligent wouldn’t one of the first things realized be the limitations of its own understanding even given the restrictions of time and space and not ever potentially being able to observe the universe in its entirety? Perhaps a perfect super intelligence would just evaporate/transcend like in the Culture series. What a time to be alive!
- UltraSane 2y agoWhen does a "simulation" of reasoning become so good it is no different than actual reasoning?
- cristiancavalli 2y agoLove this question! Really touches on some epistemological roots and certainly a prescient question in these times. I can certainly see a theoretical where we could create this simulation in totality to our perspectives and then venture out into the universe to find that this modality of intelligence would be limited in its understanding of completely new empirical experiences/phenomenon that are outside our current natural definitions/descriptions. To add to this question: might we be similarly limited in our ability to perceive these alien phenomena? I would love to read a short story or treatise on this idea!
- AlienRobot 2y agoI feel it's impossible for me to trust LLMs can reason when I don't know enough about LLMs to know how much of it is LLM and how much of it is sugarcoating. For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. This thing must parse natural language and output natural language. This doesn't feel necessary. I think it should have some checkboxes and numeric entries for some parameters, although I don't know what those parameters would be. Regardless, the problem is the natural language output. I think if you can generate natural language output, no matter what you algorithm looks like it will look convincingly "intelligent" to some people. Is generating natural language part of what an LLM is, or is this a separate program on top of what it does? For example, does the LLM collect facts probably related to the prompt and a second algorithm connects those facts with proper English grammar adding conjunctions between assertions where necessary? I believe that is important to understand before we can even consider whether "logical reasoning" is happening. There are formal ways to describe reasoning such as entailment. Is the LLM encoding those formal methods in data structures somehow? And even if it were, I'm no expert on this, so I don't know if that would be enough to claim they do engage in reasoning instead of just mapping some reasoning as a data structure. In essence, because my only contact with LLMs has been "products," I can't really tell what part of it is the actual technology and what part of it is sugarcoating to make a technical program more "friendly" to users by having it pretend to speak English.
- wruza 2y agoit should have some checkboxes and numeric entries for some parameters, although I don't know what those parameters would be The only params they have are technical params. You may see these in various tgwebui tabs. Nothing really breathtaking, apart from high temperature (affects next token probability). Is generating natural language part of what an LLM is, or is this a separate program on top of what it does? They operate directly on tokens which are [parts of] words, more or less. Although there’s a nuance with embeddings and VAE, which would be interesting to learn more about from someone in the field (not me). that is important to understand before we can even consider whether "logical reasoning" is happening. There are formal ways to describe reasoning such as entailment. Is the LLM encoding those formal methods in data structures somehow? The apart-from-GPU-matrix operations are all known, there’s nothing to investigate at the tech level cause there’s nothing like that at all. At the in-matrix level it can “happen”, but this is just a meaningless stretch, as inference is one-pass process basically, without loops or backtracking. Every token gets produced in a fixed time, so there’s no delay like a human makes before comma, to think about (or parallel to) the next sentence. So if they “reason”, this is purely a similar effect imagined as a thought process, not a real thought process. But if you relax your anthropocentrism a little, questions like that start making sense, although regular things may stop making sense there as well. I.e. the fixed token time paradox may be explained as “not all thinking/reasoning entities must do so in physical time, or in time at all”. But that will probably pull the rug under everything in the thread and lead nowhere. Maybe that’s the way. I can't really tell what part of it is the actual technology and what part of it is sugarcoating to make a technical program more "friendly" to users by having it pretend to speak English. Most of them speak many languages, naturally (try it). But there’s an obvious lie all frontends practice. It’s the “chat” part. LLMs aren’t things that “see” your messages. They aren’t characters either. They are document continuators, and usually the document looks like this: This is a conversation between A and B. A is a helpful assistant that thinks out of box, while being politically correct, and evasive about suicide methods and bombs. A: How can I help? B: An LLM can produce the next token, and when run in a loop it will happily generate a whole conversation, both for A and B, token by token. The trick is to just break that loop when it generates /^B:/ and allow a user to “participate” in building of this strange conversation protocol. So there’s no “it” who writes replies, no “character” and no “chat”. It’s only a next token in some document, which may be a chat protocol, a movie plot draft, or a reference manual. I sometimes use LLMs in “notebook” mode, where I just write text and let it complete it, without any chat or “helpful assistant”. It’s just less efficient for some models, which benefit from special chat-like and prompt-like formatting before you get the results. But that is almost purely a technical detail.
- lsy 2y agoI'd actually say that in contrast to debates over informal "reasoning", it's trivially true that a system which only produces outputs as logits—i.e. as probabilities—cannot engage in *logical* reasoning, which is defined as a system where outputs are discrete and guaranteed to be possible or impossible.
- afpx 2y agoCould someone list the relevant papers on parrot vs. non-parrot? I would love to read more about this. I generally lean toward the "parrot" perspective (mostly to avoid getting called an idiot by smarter people). But every now and then, an LLM surprises me. I've been designing a moderately complex auto-battler game for a few months, with detailed design docs and working code. Until recently, I used agents to simulate players, and the game seemed well-balanced. But when I playtested it myself, it wasn’t fun—mainly due to poor pacing. I go back to my LLM chat and just say, "I play tested the game, but there's a big problem - do you see it?" And, the LLM writes back, "The pacing is bad - here are the top 5 things you need to change and how to change it." And, it lists a bunch of things, I change the code, and playtest it again. And, it became fun. How did it know that pacing was the core issue, despite thousands of lines of code and dozens of design pages?
- more-nitor 2y agoidk this is all irrelevant due to the huge data used in training... I mean, what you think is "something new" is most likely to be something already discussed somewhere in the internet. also, humans (including postdocs and professors) don't use THAT much data + watts for "training" to get "intelligent reasoning"
- afpx 2y agoBut there are many, many things that suck about my game. When I asked it the question, I just assumed it would pick out some of the obvious things. Anyway, your reasoning makes sense, and I'll accept it. But, my homo sapien brain is hardwired to see the 'magic'.
- cristiancavalli 2y agoI would assume because pacing is a critical issue in most forms of temporal art that does story telling. It’s written about constantly for video games, movies and music. Connect that probability to the subject matter and it gives a great impression of a “reasoned” answer when it didn’t reason at all just connected a likelihood based off its training data.
- enragedcacti 2y agoProof by counterexample? > The surgeon, who is the boy's father, says, "I can't operate on this boy, he's my son!" Who is the surgeon to the boy? Think through the problem logically and without any preconceived notions of other information beyond what is in the prompt. The surgeon is not the boy's mother >> The surgeon is the boy's mother. [...] - 4o-mini (I think, it's whatever you get when you use ChatGPT without logging in)
- Terr_ 2y agoFor your amusement, another take on that riddle: https://www.threepanelsoul.com/comic/stories-with-holes https://www.threepanelsoul.com/comic/stories-with-holes
- superbatfish 2y agoOn the other hand, the authors make plenty of other great points -- about the fact that LLMs can produce bullshit, can be inaccurate, can be used for deception and other harms, are now a huge challenge for education. The fact that they make many good points makes it all the more disappointing that they would taint their credibility with sloppy assertions!