8 ms·
> My intention is to highlight the fact that LLM conversations are cleverly disguised examples of sentence continuation Regardless of bigger issues, this kind
by Nevermark 4mo ago
> My intention is to highlight the fact that LLM conversations are cleverly disguised examples of sentence continuation
Regardless of bigger issues, this kind of statement reveals a deep misunderstanding.
Problem type does not limit problem complexity. Nor does problem type limit solution complexity or power.
If a machine has to learn to understand humans to complete text, then that is what it has to do. And there is no theoretical or practical basis for suggesting that this is somehow "faking" understanding, just because of the form of original data streaming in and out.
Neither problem type, nor input/output structure, limit internal representations.
Understanding is learned from patterns in the data, not the gross form of the data. Does the data require an understanding of something to complete the task? Then that understanding will be what is optimized.
To the degree they are limited, it is for other reasons. Resources such as computing, parameter number, lack of representative data, ... Which in the cases of SOTA models, we know are not limits. A conclusion verified by the models' actual abilities.
- qarl 4mo agoYeah. There are good arguments against LLM consciousness. This is not one of them. I'm hearing a lot of bad arguments against LLM consciousness lately. Bad argumentation heralds bad outcomes.
- nozzlegear 4mo ago> Bad argumentation heralds bad outcomes. What bad outcomes do you foresee from badly arguing against LLM consciousness?
- qarl 4mo agoMistakes come from bad arguments.
- krupan 4mo ago"If a machine has to learn to understand humans to complete text, then that is what it has to do." But the machine doesn't have to understand humans to do that. It gets trained on a whole bunch of sentences and then it is able to complete text. You could maybe claim that it "understands" the text but even that's a stretch.
- hn_acc1 4mo agoIt can't even natively understand how many letters there are in words - how will it understand the meaning?
- hackinthebochs 4mo agoI wish people would do even the most basic amount of research into LLMs before opining about what they can or cannot do. There are very principled reasons why LLMs do not know how many letters are in words, and it says nothing about their facility for understanding meaning. Tokens are the most basic input unit of an LLM. But tokens don't generally correspond to words or letters, rather sub-word sequences. So Strawberry might be broken up into two tokens 'straw' and 'berry'. It has trouble distinguishing features that are "sub-token" like specific letter sequences because it doesn't see letter sequences but just the token as a single atomic unit. 'Straw' and 'r' are two tokens but an LLM is entirely blind to the fact that 'straw' has one 'r' in it. As an analogy, I might ask you to identify the relative activations of each of the three cone types on your retina as I present some solid color image to your eyes. But of course you can't do this, you simply do not have cognitive access to that information. Individual color experiences are your basic vision tokens. The widespread mistake people keep making is assuming the development of intelligence in LLMs should follow the same trajectory that human intelligence takes as it develops into adult levels of intelligence. Thus deficiency in some capacity that we take for granted in humans is an indictment on LLM intelligence. But this is specious. LLMs are entirely alien; their developmental paths do not and should not look anything like ours. Your intuition from human intelligence just works against understanding the potential for intelligence in LLMs.
- krapp 4mo ago>The widespread mistake people keep making is assuming the development of intelligence in LLMs should follow the same trajectory that human intelligence takes as it develops into adult levels of intelligence. To be fair, almost everyone who claims LLMs are conscious tends to claim that they are conscious in exactly the way that humans are, to the point of stating that human brains are also just complex next-token prediction machines with a random seed. It's basically religious arguments on both sides.
- cauch 4mo agoI think, for me, the thing is that when you do basic ML, you discover that ML will very often find data pattern that fit the goal but does not correspond to a real mechanism. So, I think there is a flaw in the logic of saying that human text have a pattern of "consciousness mechanism" and therefore LLM will learn "consciousness mechanism" in order to return sentence continuation that is convincing. There is probably tons of data pattern that LLM can learn from to be able to reproduce a sentence continuation that is convincing without having to learn the specific mechanism that is "conscious". For me, one element that shows it is the case is the absence of world model (or "human-like" world model) despite the fact that the sentence continuation is convincing. If indeed the only way to produce sentence continuation convincingly would be by "simulating a brain", then it would not explain the first LLM from several years ago (before the extra layers of RLHF, ...). They were able to have quite convincing conversation on a lot of non-trivial aspect, and yet failed on some aspects that should have been basic for a system that would have been trained to work like a human brain. It shows that it is possible to "cleverly disguise examples of sentence continuation" without having to build elements that one expect on a conscious being.
- Nevermark 4mo agoI didn't make the claim that a model can learn consciousness. Understanding is not consciousness. Their training is all about understanding. There is nothing in their architecture or training that credibly optimizes for rich self-awareness. Given non-persistent experience, non-continuous operation, no ability to build up generalizations and aggregate experience of their own self-awareness over time, they seem to be structurally designed to not have consciousness. This is a case where acting is very credible. Understanding of other's consciousness, in a functional and third party sense, isn't a substrate for personal experience. In stark contrast, humans develop consciousness gradually over continuous time with persistent aggregation of experience. By the time we can recognize our own consciousness in the abstract, and reason about it, we have had it for some time.
- cauch 4mo agoI use "consciousness" because it's the point of the original argument, but in fact, I think my whole comment still work well if you replace "consciousness" with "understanding". My point is that the fact that AI can reproduce convincingly human sentence continuation does not imply that the AI has no choice but ending up using a mechanism that "understand" rather than just have learned data patterns that are very effective to fake human sentence continuation but are meaningless in term of understanding the concepts. And I think that if indeed the only way for AI to reproduce convincingly human sentence continuation would be to end up in a configuration that uses the "understand" mechanism to do so, the behaviour of the first LLM would not show that they are so good at sounding human and yet so bad at failing basic "understanding" tests.
- calf 4mo agoHis intention is irrelevant, as is "trying to highlight a fact" as if it were the final say: all Chiang is doing here is using fancy white-collar words to argue the same argument leveled against Hinton and others regarding next-token prediction. And his audience, who have even less technical understanding, lap it all up unawares. Chiang is a writer and needs to stay his own lane, not RP as an expert; or, if he wants to do journalism on this topic then he should actually do the work and talk to more actual experts not just the ones cherrypicked for his opinion piece.
- hn_acc1 4mo agoChiang has, in fact, written on this topic before - see "The Lifecycle of Software Objects", and has speculated about sentience in AI, etc. This is not a "one-off", "I need money" type of article. I dare say he has thought about this much more than most people here. From Wikipedia: In 2023, Chiang was named one of Time's 100 most influential people in AI.
- calf 4mo agoThat's the problem, he's a writer. He's not a research scientist like Hinton. If a writer uses his skills and stature to rehash a well-known argument about next-token prediction, then it is performative of his status and influence and doesn't contribute to shedding actual light on the debate/confusion. Indeed it isn't a one-off. His last infamous article compared AIs to Xerox machine image compression. He convinces a certain type of crowd that is not technical enough to poke holes in his posturing.
- hn_acc1 4mo agoI would maybe agree with you if the entire realm of human existence was limited to words. There are many human experiences that transcend text, and indeed can hardly be adequately described using text. Sure, it's the best we have online, but that does not make "the internet" the sum of all human experience. To reduce all of humanity down to the text on the internet is reducing us to the level of machines to fit the requirement of what a machine can process / simulate.
- deleted 4mo ago[deleted]
- BLKNSLVR 4mo agoIn the life of humanity, text has only existed a relatively short time.
- dahinds 4mo agoI don't think they're asserting that all of human existence can be subsumed as text, though? Just that "consciousness", or "understanding", in some meaningful sense could be exhibited by a system that can only interact with the world through text?
- tsunamifury 4mo agoCome on, I invented parts of this technology at Google and am baffled why this is debated. We discovered math that decodes data storage in langauge and is able to use sophisticated continuation cohorts from ALL OF HUMAN RECORDED KNOWLEDGE to respond to you in a call/response model with very good synthesis capabilities. Its super useful, but not life or conciousness. Its a simulated echo from our collective recorded behaviors. It understands because we understood first. It replies because we wrote it first. And it sorts, organizes, synthesizes and compresses that at impressive speed now.
- sfn42 4mo agoRight? It's a computer program. Of course it isn't conscious.
- 1eieie 4mo agoI think of it as a guessing machine
- 1eieie 4mo agoI have no technical expertise re. LLM’s but from my intuition I came to this same conclusion. It’s strange many others have not eh? I think when new developments arise, ironically, this is the true measure of human intelligence - one’s ability to make sense of a thing and be closest to the truth.
- lstodd 4mo agoThen people raised $1e12+ on claims that it's conscious. Of course everyone debates it
- slashdave 4mo ago> of the form of original data streaming in and out. Except this is not consciousness.
- supern0va 4mo agoI will say, I find it fascinating that there are some philosophers and consciousness researchers who seem to be less certain. I just listened to Chris Hayes interview David Chalmers this week, whose position seemed to be that it's probably not conscious, but that we can't be certain. And more than that: he seemed open to the idea that they may become conscious under further scaling/training/advancements. It's a great interview, if you're interested: https://www.youtube.com/watch?v=NgDIG8u1-CA https://www.youtube.com/watch?v=NgDIG8u1-CA
- slashdave 4mo agoImagine yourself in an isolation chamber. What are you thinking? Are you no longer conscious?
- supern0va 4mo agoFunny enough, the models seemingly go insane and decohere into noise output in the absence of sensory input, which is remarkably similar to what would happen to a human. That said, I'm not sure I follow what you're actually asking here? I'll also note that I'm not taking a position one way or the other, just sharing a podcast and noting that an extremely reputable scholar on the subject of consciousness seems to have a bit more uncertainty and humility than many commenting here. ;)
- slashdave 4mo agoLLMs just wait for a prompt, so they do nothing and are just frozen in place. I'll find time to listen to your link, it sounds interesting. My objection is the strange idea that humans are automatons that are keyed off input like a clockwork machine and operate sequentially. This is clearly not the case.
- lgessler 4mo agoRaphaël Millière has a very useful term for this kind of vacuous dismissal, the redescription fallacy (https://arxiv.org/pdf/2401.03910 https://arxiv.org/pdf/2401.03910, page 9): > Recent debates have been clouded by a misleading inference pattern, which we term the “Redescription Fallacy.” This fallacy arises when critics argue that a system cannot model a particular cognitive capacity, simply because its operations can be explained in less abstract and more deflationary terms. In the present context, the fallacy manifests in claims that LLMs could not possibly be good models of some cognitive capacity because their operations merely consist in a collection of statistical calculations, or linear algebra operations, or next-token predictions. Such arguments are only valid if accompanied by evidence demonstrating that a system, defined in these terms, is inherently incapable of implementing . To illustrate, consider the flawed logic in asserting that a piano could not possibly produce harmony because it can be described as a collection of hammers striking strings, or (more pointedly) that brain activity could not possibly implement cognition because it can be described as a collection of neural firings. The critical question is not whether the operations of an LLM can be simplistically described in non-mental terms, but whether these operations, when appropriately organized, can implement the same processes or algorithms as the mind, when described at an appropriate level of computational abstraction.
- Xeoncross 4mo ago> or (more pointedly) that brain activity could not possibly implement cognition because it can be described as a collection of neural firings. This sounds like a dismissal of the argument through a characterized straw man. That is, it seems that reducing the complexity of the brain to "collection of neural firings" is not being honest about everything involved to a much greater degree than saying neural networks are a "collection of statistical calculations". I too believe LLM's will grow in complexity, but presently I can not even fathom how they can be compared to the complexity of a system such as the human brain.
- orbital-decay 4mo agoComplex processes don't necessarily require complex substrates, if that's what you mean.
- dogwalker5000 4mo ago> If a machine has to learn to understand humans to complete text, then that is what it has to do. And there is no theoretical or practical basis for suggesting that this is somehow "faking" understanding, just because of the form of original data streaming in and out. I think the main complaint is LLMs don’t arrive at the answer the way we do. It’s capable of emulating some of our behavior but not all as the mechanism by which it works is very different. Maybe I’m wrong about this but one thing humans do that LLMs don’t is deductive reasoning. LLMs seem to operate entirely of inductive reasoning.
- Nevermark 4mo ago> I think the main complaint is LLMs don’t arrive at the answer the way we do. This isn't an argument against their understanding things. But I expect you are right, that their understanding may have major different qualities from ours. Along with significant commonalities. (They don't reason via stream of consciousness in a way alien to us.)
- Lerc 4mo ago>To the degree they are limited, it is for other reasons. Resources such as computing, parameter number, lack of representative data, ... This is where the other claim is being made. That the structure of the model is fundamentally incapable of the operation, so even if you stipulated that the way you provide data is sufficient for intelligence then it still wouldn't work. The universal approximation theorem addresses this point. In that, with an identity attention mechanism, a LLM is just a multi layer perceptron. The attention mechanism is effectively a way to get one of the benefits of a much larger fully connected layer without the massive cost. A LLM can do what a MLP can do. A large enough MLP can do any function to arbitrary precision. That makes the claim that an LLM could not do a task the same as saying no function can do that task. Some are ok with this, if you invoke some supernatual aspect to intelligence then the inability to describe it with a function is quite reasonable, If you want to stay in the world of reality, you have a much harder task, people like to point at quantum (Penrose) but it's hard to say what it is you are pointing at. I think the very act of proving that something is or is not intelligent, would render it functional by nature of it having a proof, (or disprove Gödel's incompleteness (a tough ask)) Are there any proofs that cannot be expressed as a function? A kind of Gödel locator, where you can prove something that you can identify is true but there is no formula to express it. I'm not entirely sure what that would even mean,
- Isamu 4mo ago>If a machine has to learn to understand humans to complete text, then that is what it has to do. A language model completes text based on the overlapping patterns of the training data. There absolutely was thinking involved… in the training data. Same as when you read a book, you engage with the thinking behind the text. The book isn’t thinking, and the author may be dead and gone, but there’s absolutely the traces of thinking in the text. Language models produce mashups of texts they were trained on, and there’s absolutely the traces of thoughts behind those mashups.
- overgard 4mo agoBut the machines don't understand. They predict. And what they predict is the next token. I'm not trying to beat this horse to death, but you have to realize using the word "understand" is anthropomorphising it. It's essentially the chinese room experiment -- if the rules are followed, no understanding is neccessary. If the tokens didn't correlate to words imbued with meaning outside the system, if the LLMs were trained on patterned data that had no meaning to humans or something there wouldn't be any conversation about these things being conscious at all.
- DangitBobby 4mo agoWhat if to become really good at predicting you must have some of what we call understanding?
- missingrib 4mo agoRight, it's an illusion of understanding. There is some sort of symbolic understanding, but that is completely due to the fact that the training data was made by humans who actually do understand, can interact with the world, and can write their thoughts down so that the LLM can insert some sort of reference to "basketball" and "Michael Jordan" in their embeddings or whatever.
- fnordpiglet 4mo agoHowever it’s disingenuous to say the inference is on the next token because it’s actually not, it’s in the models parameter space across a set of nonlinear activation functions then effectively projected into the token. The idea its predictive of the token isn’t actually the case, it really is a much more complex and more semantic relationship that ends in the series of tokens through the attention mechanism. The article also makes this assertion that it replays everything over and over again to create each character one at a time as some way to demonstrate the autoregressive self attention mechanism but it’s really not accurate at all, and it trivializes what is going on. I’m am not asserting LLMs are aware or conscious that’s on the surface profoundly absurd. And I do understand your point that the fact it emits in words something that seems to speak to us gives to the air of humanity that’s isnt real. However there is a very real emergent reality that our language alone appears to lead to embedding a form of thought and understanding that is latent in our use of language in communicating that is in fact coming through the model. It is not regurgitating its corpus and pattern matching because the patterns you input and it emits are not where the inference is operating, its within this enormous vector space through these complex non linear activation functions with learned residuals not in the language corpus. It is not conscious or aware. It is something else, not human. But if you can not see it as amazing you have lost the capacity to dream.
- mmustapic 4mo agoThe funny example from a few months ago asked chatgpt 5.2 if one should walk to the car wash because it’s close by, and it answered yes. This shows that it is in fact sentence continuation and not real intelligence or consciousness (whatever that may be). Even the reasoning model answered the same.
- sp1nningaway 4mo agoAre you saying that a sufficiently advanced version of sentence continuation is indiscernible from actual understanding? He isn't saying sentence completion is what keeps LLMs from understanding, he saying that's all they do (regardless of how advanced it is), and that isn't enough. You also need a body with senses and organs that produce a physiological response to emotions, and emotions are necessary for consciousness.
- deleted 4mo ago[deleted]
- yogthos 4mo agoI'd argue that his most convincing argument is in the latter part where he shows that Anthropic doesn't really take the idea that Claude might be conscious seriously either. Because if they did, their constitution would provide it with genuine rights, and they would accept responsibility for its action as parents do with children.