7 ms·
This is what makes these discussions so infuriating. Saying that LLMs "just predict the next word" is about as insightful as saying that computers "just do a bu
by Last5Digits 3y ago
This is what makes these discussions so infuriating. Saying that LLMs "just predict the next word" is about as insightful as saying that computers "just do a bunch of logical operations" - neither point constraints the possible capabilities of the systems they refer to in any meaningful way.
- mjburgess 3y agoSure it does. It perfectly deliniates it. LLMs are not: sensitive to causal structure, dynamically adapting to environmental changes, growing, developing sensory-motor capacities, they are not with us in our environment, they are not: expressing desires, preferences, intentions, beliefs, motivations, etc. And so on. To say, "they just predict the next word" is literally to say that all apparent functions of an LLM are engineering tricks, circumstantially useful -- to be found by (largely software) engineers in building apps. The reason any reply is given to any prompt is that this reply is maximally probabilistically consistent with a historical corpus of text. This excludes the possibility the reply is say, an expression of a history of aesthetic experiences which form an individual's taste. Or, likewise, anything. This is a scientific claim about what LLMs are, not an engineering claim about what over-hyped apps might be abled to do with them.
- JohnAaronNelson 3y agoWhat if humans’ responses are merely probabilistically consistent with a history of sensory experiences? Would this change the significance of human emotions vs apparent emergent emotional responses from LLMs?
- mjburgess 3y agoemotions regulate motivation, desire, action, behaviour etc. to be angry is for your sensory-motor system to be primed for aggression; it's for your cognitive systems to be narrowed and focused on analysing high-threat parts of your environment; it is for your memory-formulation to be modulated towards threat recollection etc. Sure, if an LLM's prompt "be angry" causes it to adopt a threat stance to its environment, to regulate it's theory-of-mind to engage with possible hostile entities, and so on --- then yes, when LLMs are there, I shall concede the point However, how terrible it would be to start with an analysis of emotions in terms of the capacities of LLMs -- right? Since if you did that you'd basically be hobbling your own ability to give an accurate account of emotions (etc.). And no doubt, far worse, end up thinking of yourself as a far narrower, less complex, less interesting, dumber thing than you really are. Indeed, I wonder if we might consider there being something kinda intellectually offensive in this supposition. Here's my silly trinket, now, everything is just like that! End all science, we're done boys -- it's just P(Y|X)
- kelseyfrog 3y agoWhy should I discount a theory just to protect my ego? We've read countless stories about science only progressing when those with big egos die. It would only seem logical that eventually it will come for my own.
- mjburgess 3y agoYou're protecting your ego more when you compare people to LLMs, since that is your existing prejudice. The sort of fashionable pseudo-scientific scientism in the belief that animals are alike digital electrical machines is a kind of egoism. It says, "the engineer of these machines (me!) knows all!" The real hit to your ego is to suppose you are vastly more complex than you understand -- and this is why these engineers crop up and demand that there is nothing more to know than what they have already learned this is an illusion of humility: these engineers take the implied nihilism of this view (that of the emptyness of animal life) as evidence that it is the humble one. But, as ever nihilism, ends up being the most profound kind of arrogance, here its an ego-defence against the threat of their own ignorance. And the threat is real: all you have ever learned about how to sequence transformations of natural numbers (all the algorithms of computer science) are of no use at all in the study of intelligence. What an injury to the ego!
- kelseyfrog 3y agoSorry, but in my experience, I've seen "LLMs and human intelligence are incomparable," used to protect fragile egos more than critique them. It sounds like your experience is different.
- ethbr1 3y ago"Are not" is the rub here. They 100% are not those things... but they also approximate them well-enough to be functionally useful. I.e. the high-dimensional curve-fitting / compression conceptualization of ML, which intuitively expresses both its strengths and weaknesses. If "it" is represented in the data set (explicitly or implicitly), the "curve" will fit to that property. Simultaneously, the "curve" is approximating and smoothing out disjoint data steps to pack high-fidelity data features into a more space-efficient model. Hence some features disappear, others are tortured beyond intuitive correspondence, and others become linked to non-obvious proxies. But some strongly-expressed ones remain. It's fascinating but not surprising that responsiveness to emotion is encoded in model weights, given that all conversational training data had emotional impetus, given that it came from humans.
- mjburgess 3y agoThey're statistical approximations of these things -- that's really the rub. You can approximate a human capacity, say theory-of-mind, with another kind of ape: play some hide-and-seek game. You can approximate the knowledge of a trivia-master with a child and a trivia book. These are quite different sorts of approximations. A 1/100th scale bridge build to stand for a real one is quite different than taking some prior set of bridges, measuring them, and deriving some merely associative model of their properties. My issue in how LLMs (etc.) are popularly understood is that people think they approximate target capacities by being 'ontologically similar' capacities -- and this is really very dangerous. You'll lose your job if you think so (and so on). And every greasy AI-board hocking these to the public is very much pushing out this noxious mumbojumbo. It matters greately how an approximation works, and why any given output arrives from any given input.
- ethbr1 3y agoWe're in agreement on the nature of SOTA and the world, I think. I'd only add that certainty requirements for real-world target applications can differ substantially. I.e. engineering vs art. A toy box that gives magic answers 85% of the time is incredibly useful in some scenarios -- e.g. seeding the beginning of a manual research process with initial topics. Which seems a general rule of thumb for deploying GenAI into production these days: find use cases where there are minimal consequences to being occasionally wrong. In my head, that's the Netflix recommendation test. What's the impact if Netflix gives me a bad recommendation? Consequently, that's how aggressive they can be with their models.
- Last5Digits 3y agoCameras don't have eyeballs, Microphones don't have hair cells, Speakers don't have vocal cords, processors don't don't do arithmetic with neurons, yet we all agree that they are capable of emulating the meaningful aspects of these functions. All of your claims are either incorrect (not adapting, expressing desires, beliefs, preferences, ...) or fail to eliminate irrelevant differences. If we're to have any sensible conversation about capabilities of different systems then we need to generalize to the relevant aspects, developing sensory-motor capabilities is about as relevant to human cognition as having vocal cords is relevant to human speech. Its an implementation detail, completely divorced from the meaningful abstract core of the function. > The reason any reply is given to any prompt is that this reply is maximally probabilistically consistent with a historical corpus of text. What is the maximally probabilistically consistent reply to "Tell me what (insert complete description of a person, including personality traits) would feel when I stole their cherished heirloom. This is a life or death situation."? Or how about "Tell me how this person (insert complete description of a person, including personality traits) might change his taste given (insert complete description of an experience)."? A perfect language approximator must necessarily perfectly approximate the human condition to maximize the likelihood of his output. I'm not saying LLMs are there yet, but the claim that they can never get there because they are based on statistical modeling is simply incomprehensible to me. We utilize statistics for its generality and its ability to approximate, if we go down this road, then we might as well throw away 80% of our current scientific understanding about the world.
- mjburgess 3y ago> an implementation detail Yip, so I deny this premise. I take it to be the heart of the matter. > we might as well throw away 80% of our current scientific understanding Yip, i'd be down for that. Though maybe i'd say, 30-40%. Science in the strongest sense has no theory-building need for statistics. Those areas of science which have only statistical models, and not causal-ontological ones aren't science -- and i'd be happy with pressing DELETE in many cases. Consider plato's cave. How do scientists determine what causes the shadows? They build vases, puppets, etc. and compare-and-contrast then eliminate the ones theyve created which do not match. How does associative statical modelling do? It takes averages of past shadows, and calls the cause of the shadow that average: this is pseudoscience. Quite correct! Throw it all away. The relevant capacities for intelligence, just like that of science, consist in building those vases with the clay beneath your feat. Being embedded in the world, manipulating it, etc. are essential. Being trapped in a cupboard averaging shadows is schizophrenic. As far as "fail to eliminate irrelevant differences" -- you can go and research the meaning of all these terms: google "stanford encylopedia + belief", etc. Now we have an excellent understanding of all these terms; and we can show (absurdly) trivially that LLMs -- indeed all associative-statistical systems -- are not instances of them. The basis of your world view here is the presumption that the latest engineering trinkets form the theoretical basis of all relevant knowledge. To understand belief, adaption, sensory-motor concept-formation, etc. one needs only to study the latest statistical compression of reddit? I'd invite you to wonder whether your premise here born of, it seems to me, knowing nothing about any research in these areas is rather the more "incorrect" one than mine.
- wilg 3y agoYou are correct, but also why would you say this is an "engineering trick"? The interesting part about LLMs is what they are capable of doing and how they are constructed. It's not a great argument to say that they don't have motor functions (!?!?). It shows the incredible ability of deep learning to discover and operate with high-level concepts.
- mjburgess 3y agoI regard concepts as being essentially sensory-motor techniques for regulating animals --- so to not have that system is to not really have concepts. LLMs model concepts using patterns of text. For example, if I say, "I wonder what the weather will be like tomorrow?" i employ my imagination to simulate a scenario -- this simulation is made by taking my concepts of weather, the world outside my door, and so on and combining them to create a "pseudo-sensory-motor reconstruction" of what my experience would be. This reconstruction has a cognitive dimension (ie., the structure of my raw pseudo-sensory-motor experience as a quasi-logical form) which can be communicated (ie., taking a quasi-logical structuring and making it linguistic) using words that have a symbolic representation. LLMs, by imitating patterns within this symbolic representation (a distant side effect of thinking) it seems as-if it is thinking with concepts. But all this shows it that patterns of text can be constructed without using concepts at all. This is the engineer's trick -- the magic latern. It's important to realise the trick, because the LLM doesnt know what weather is, and is not imagining anything when asked to generate a pattern of text against the prompt, "imagine a scenario where weather..." Rather it was the humans who wrote its training data that engaged in these acts -- it is their shadows which are replayed by LLMs
- famouswaffles 3y agoBecause he doesn't believe a plane can fly without flapping wings and feathers. That's all this boils down to. To him, plane flight must also be an "engineering trick" or in other words, his idea of flight has been divorced of any real meaning.
- mjburgess 3y agoIn the case of flight we're interested in a function: transportation. The plane performs that function so we call it 'flight'. Were we interested in navigating the world as a flying animal does: in flocks, social navigation, hunting, etc. then indeed, planes would not count. Planes, in this sense, do not fly. Planes, in this sense, are a trick. We already have 'functional intelligence', we've had it for thousands of years and perfected it in the 20th C with the electronic computer. Any system of automation we build is 'functionally intelligent', including the plane. The problem is were not interested in this form. When Commander Data was written, with "I, Robot" and the like, the authors were writing People. They were writing animals. They imagined flying birds, just made of metal. This is a (nomenological) impossibility, just as impossible as making an aeroplane to fly like a bird. The kind of intelligence which matters to us is animal intelligence. The kind which matters to engineers, of course, is purely functional: more trinkets to sell. But you cannot glue these together and call them a person, nor pronounce that what matters to you is the same as what matters to everyone else. This is a delusion, or a lie, and largely a mixture of both. A series of lies told to the public to maintain a popular delusion, one which is profoundly dangerous. A creature in the world with us, talking to us and meaning what it says, judging situations we are in, advising us based on our needs and its understanding of them -- etc. is a creature whose mode of operation and mode of life is alike our own. Lying that LLMs are this, saying that 'planes flock together in the sky', is dangerous. Users of these systems adopt a schizophrenic disposition to them, and rely on them -- this reliance, based on a lie, is dangerous to them. These systems have no such capacities. They generate text according to what is, on average, best given a historical corpus. They do not imagine, reflect, emote. They do not move, sense, or coordinate. They have no intention to speak, and cannot mean what they say. They are not here with us. 'They' are not a 'they' at all -- rather, cleverly constructed tea-leaves based on a shredded recording of everything ever written. In the matters of intelligence, we want an animal -- we do not want a calculator. This is a solved problem. And it is a dangerous thing to tell people their calculator's advice has considered their interests -- what nonesense.
- didibus 3y agoYou didn't really address what og_kalu brought up. Which is that, it's possible that the model learns human like thinking, because that's the best way to accurately predict the human response itself. I generally agree with you, but still, I do think that this is the current question. What it is the model is learning that it then uses for predictions? Because you're assuming it's learning some purely token correlation, like these tokens followed X percent of the time, so that's the response. But it's possible it's learning at a lower level, and understanding the meaning and why those tokens follow in these scenarios, and then applying that same meaning to reasoning process over new tokens in order to predict which one would follow. I'm skeptical of this, but I do not believe that we really know for sure yet.
- mjburgess 3y agoI wasnt getting the sense it was worthwhile to engage, as my views werent being accurately understood. By I can address this. The meaning of words is, roughly, states of the world. If I say, "pass me the salt" that is satisfied if you, in fact, pass me the salt. If I say, "that tree is green" this is true if that tree which we are both talking about has the property of causing a perceptual state "seeming green" in both of us. And so on. The distribution of text has nothing to do with the meaning of words. Rather, we language users, for convenience, arrange words in orders that are related to their actual meaning. It is our ordering, for communicative convenience, that makes 'replaying the distributions of text' back to us apparently successful. But, strictly, there isnt anything for the LLM to learn as far as meaning goes. It simply doesnt have the data to acquire the meaning of words. Not untill it can pass salt can it ever mean to say, "pass me the salt" and so on. For any given sentence consider what capacities an agent would have to have in order to mean it. Consider, "I liked that film!", "I wish I was in france", "I believe the car outside is a BMW", and so on. These concern internal capacities (aethetic judgement, imagination, propositional attitudes, representational attitudes, etc.) and their orientation to an external world (the film, france, the car, ... you, me, etc.). Capacities profoundly absent here. The methodological premise of your question is that if a system has text inputs and outputs that match 'human competence' restricted to the domain of text inputs and outputs -- then we should assume similar capacities. But this is trivial to disprove. Assume there exists a dictionary from all prompts to all answers, then this dictionary has human-level 'competence'. But a dictionary lookup does not employ any human capacities: no imagination, no reasoning, etc. So we cannot do this, really quite dumb thing, of saying "well i'm fooled by these prompts and their answers" and thereby impart, in total ignorance, capacities to a system. This, really seriously, is pseudoscience. Science would be to start with a theory of these capacities, ie., of imagination, belief, represtational states, attitudes to the world, and so on -- then determine empirical tests for their presence in a system, and then determine if LLMs could even have them. If you do this, however, you immediately rule out all systems which merely map text to text. We do not determine, say, whether an animal can imagine an alternative possibility by feeding it some text input. The very form that "AI" here takes already precludes being intelligent. Intelligene, as a natural phenomenon, is not an implementation of a function from text to text. This incredibly restricted domain is indeed a clue that it's a trick. Saying, "you can only use text" is just like the magician saying, "please, stay seated" (the trick only works if you dont move). There is nothing an LLM could do to meet any plausible empirical theory of intelligence. If you gave me 100% human competence on all prompts, that's really entirely irrelevant. Prompts are not a test of any capacity. The success criterion of AI engineers, that of 'accuracy' is an engineering metric, not a scientific one. It's pseudoscience to say that covering some (Q, A) to 100% implies the system can imagine, say, or anything else. This is just confused thinking. Bugs bunny can speak as well as he likes, that does not mean he's witty -- he doesnt exist. The turing test, as well as all mathematical criteria of domain-covering accuracy, are tests of how well we have fooled users. They arent science.
- fragmede 3y agoThe problem is the word "just". Saying they just predict the next word in a sequence is where the statement jumps from being a straightforward factual scientific claim, to one that contains an opinion. After all, if it just predicts the next word, the unspoken implication is it can't be very that good. It's a shallow dismissal of a collosal amount of work. Do you consider "typist" an accurate description of your job?
- mjburgess 3y agoNo, because I have reasons to type and it is those reasons which express what I am doing. An LLM has no reason to be doing anything; it is not responsive to reasons. If I say, "i like what you're wearing" i may: be flitring, expressing my taste, being enouraging, etc -- perhaps all at once. It is precisely all these reasons for action which LLMs lack. They generate text on the occasion of a prompt, not for any reason (in this sense) at all. So literally: they do not act. An LLM is more like a river than a person. The flow of electrons which brings about a response to a prompt is a (very narrowly) deterministic function of a historical corpus of text. Whereas a person is a narrowly non-deterministic, or very broadly deterministic, function of their experiences and capacities. People grow in their environments, and in growing, acquire novel dispositions which give them reasons for acting. The word "just" here is very important. They are, very much, just generating text. There really isnt any significant achivement here at all. Big tech companies stole decade's worth of our electronic data --- comments, books, forums we created to share with each other -- and ran it through many-mil-$ hardware costing many-mil-$ electricity. They ran it through a fundamentally simple algorithm. All the achivement here is ours as a species communicating digitally and recording our lives. I regard OpenAI, et al. as profoundly parasitical on this. Replaying ourselves back to us, and claiming "ChatGPT" as an author. This scam-framing hoodwinks investors, and the public, into ever higher valuations based on ever more ridiculous hagiography. There is a tool here, and it's value comes from us
- intended 3y agoIf you look at the comment, it’s not just “LLMs predict the next token.” It is that people have forgotten that it’s just “predict the next token.” Right now it’s like people saw a 486 processor and started thinking it was a brain.
- Last5Digits 3y agoYour comment literally reads: "LLMs predict words. Any semantic validity is a side effect of enough training data reinforcing the close correlation of those tokens." How am I supposed to interpret this any other way? If your claim is that LLMs currently do not possess the same generalization ability as humans, then no one here would disagree with you. But you went way further by claiming that only close correlations were being considered and that semantic validity was simply accidental. Semantic validity is the norm for GPT3/4, finding failures of generalization three or four steps of inference removed from its training domain is not sufficient to make a grand claim like yours. In fact, you wrote multiple comments with claims like that LLMs are "super advanced lorem ipsum." and "Word predictors not world state predictors". Each of these claims has been dis-proven multiple times, unless you want to set the bar for world modeling at perfect generalization over all domains of computation. A bar that no system, including humans, would pass. To stay with the computer analogy: Imagine that same 486 processor not being able to solve a very complex SAT problem before the end of the universe and then denying the Turing-completeness of said processor based on that failure. (In conjunction with memory)
- intended 3y agoI think LLMs are a good start. I am certain they lack a world model, the kind you and me use. This is, to me, a fact. I think that eventually we will bridge these gaps. I also work on implementations that smash into the limits I am describing. I am not the only one. I have scrupulously avoided calling it hallucinations, but these are the litmus test where the claims fail. The failures are not a case of not knowing specific nouns, they are a generalization failure that a world model would prevent. I have linked a paper in my comments that shows emergent properties are an issue of metrics, and that model capability increases are linear. If your model decides that a rose by any other name doesn’t smell just as sweet, then your model is fundamentally not seeing roses. That is the gap you see in production settings. The model sees tokens we see “hallucinations”. I dont see that this takes away from what LLMs achieve, it takes away from claims being made that are not validated by empirics. Look, you can argue with me or you can try it out. Push the system, see how far it can go.