6 ms·
Well LLMs don't learn anything. They don't know anything. And they certainly don't understand anything.
by Denote6737 3y ago
Well LLMs don't learn anything. They don't know anything. And they certainly don't understand anything.
- golol 3y agoExcept they can play chess better than 90% of humans... Just by luck of course..No understanding.
- pxmpxm 3y agoI've got a disposable gift calculator in front of me that can do division better than every human ever. No one seems to anthropomorphize it, however.
- golol 3y agoIt has learned an algorithm that plays chess purely by reading chess logs. It can play at different skill levels depending on the surrounding natural language How is that not understanding? By the way, would you say that Stockfish 15 understands how to play chess well?
- xrisk 3y agoI’m pretty sure you can write a simple tree search based program that plays better than 90% of humans, sheerly by crunching numbers quickly. I’m sure you wouldn’t claim that that program “understands” chess in the way a human does. Of course it’s a different question as to how exactly do humans understand chess. I personally don’t believe humans have some sort of “special” intelligence, whatever we do must be computable and therefore doable by a computer. It’s just that todays LLMs don’t seem to be it.
- johnthewise 3y agoSure, that wouldn't be impressive if you wrote tree search for chess. However, if you wrote a program that doesn't do any search and still plays well, that'd be impressive. I don't understand your claim that LLMs playing chess is anything sort of understanding. related: https://parrotchess.com/ https://parrotchess.com/
- xrisk 3y agoHow do you know there’s no search involved?
- Jensson 3y agoIt does more illegal moves the more you deviate from standard chess openings and positions and becomes unable to play at some point. If it did a tree search based on chess rules that wouldn't happen. If it encoded a lot of chess openings and some tiny ad hoc logic around those then that is exactly what we would expect.
- golol 3y agoThis is not true anymore. It has been discovered that gpt-3.5-instruct when prompeted with purely PGN notation.playes at >1700 Elo and makes no/little illegal moves.
- eloisant 3y agoThey're only as good as the players that played the games they're trained on. Meaning that for computer programs they're very bad at chess.
- golol 3y agoThat is extremely impressive for a model that can do literally millions of tasks. It is not trained to play chess. It can write stories, code, recipes, franslate text, create JSON AND play chess. Don't you see how amazing this is compared to everything we had previously? LLMs are not just generalists that can't do anything specific properly. They can play chess better than most humans and I'm pretty sure soon they will play better than any human. And this in a general AI system!
- lewhoo 3y ago> Except they can play chess better than 90% of humans... Just by luck of course..No understanding. I've been thinking about this and I think there is a simple trick to make LLM's play chess well. Any chess position can be represented by a string, like so: rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1 (Forsyth–Edwards Notation) If we think of a chess game as a sentence or a sequence of such strings or tokens, then after feeding tons of those to the model it is reasonable to suspect it will do what it does best, that is predict the next token. There is no spacial awareness or anything beyond that going on I think.
- xrisk 3y agoThere are far too many possible positions aka tokens for this to be feasible.
- lewhoo 3y agoSame could be said about language but here we are. There are far more possible combinations of words in any given language than there are positions in chess for a typical length of a sentence I suspect.
- xrisk 3y agoDefinitely, but treating fen positions as tokens isn’t going to work, you need something smaller, something you can decompose into. What that is I don’t know.
- lewhoo 3y agoWell, every character of the string can be a token. From a transformer's perspective it shouldn't matter if it is Chinese, English or "chess" language.
- johnthewise 3y agoYou are right, but evidence of LLMs working well with language is not a sign that they might be memorizing Chess. But rather opposite, infeasability of memorization of chess and yet LLMs still doing well on them might be a sign that LLMs could be doing the same 'thing' with language.
- EddTheSDET 3y agoIf they’re trained on a database of hundreds if thousands of top ranked games they’ll beat 90% by following those by rote
- golol 3y agoJust to make it clear: the database doesn't have to contain exclusively top rated games. It can be 99% low ELO games, as long as there are suffiiciently many high ELO games (in absolute number), the LLM is large enough, and there is sufficient context for the model to distinguish high ELO from low ELO games, it can learn to play well given fhe right prompting.
- throwaway02y 3y ago[flagged]
- tormeh 3y agoIt’s because all the verbs are ill-defined in the context of AI. All it boils down to is “I don’t like LLMs”. You could say that LLMs can know, learn, and understand, and you’d be just as correct, just with opposite vibes.
- TZubiri 3y agoThey do learn by reading their corpus, and they do know what's in their corpus. If you created a temperature sensor and the AI removes itself from hot temps, then it feels pain. The insistency that AI not be antropormorphized sometimes gets in the way of communication.
- dbspin 3y ago> it feels pain No... It detects temperature and reacts. It does not feel pain. Feels implies both affect and sensation, both of which are experiential, phenomenological qualia requiring some corollary of a nervous system and sentience, neither of which exist within the underlying structure of an LLM or indeed any current AI. AI can no more feel pain than a sliding door can decide to open. It's deterministically reacting to stimuli.
- TZubiri 3y agoDawkins faces this argument in "The selfish gene" and he choses to explicitly acknowledge that qualia is unprovable, so takes a definition of the phenomenon that is independent of it. And this is when studying actual living organisms, so applying it to machines will be even more useful.
- throwaway02y 3y ago[flagged]
- belter 3y ago
- Alifatisk 3y agoAs Stallman said, "... it is important to realize that ChatGPT is not artificial intelligence. It has no intelligence; it doesn't know anything and doesn't understand anything. It plays games with words to make plausible-sounding English text, but any statements made in it are liable to be false. It can't avoid that because it doesn't know what the words mean."
- K0balt 3y agoThis gets into a very fascinating semantic corner. While LLMs are “merely” token predictors, it seems that perhaps the “knowing” is actually imbedded in the large corpus of memetic data. In the same way that an infinitely detailed “choose your own adventure” book shows that data and computation are interchangeable by degree, (calculation vs lookup), my conjecture is that the often surprising effectiveness of LLMs, as well as there very human like flaws, are the result of this kind of computation-in-the-data scenario. I posit that the overt connections in textual human knowledge also carry the implied knowledge that provides the unwritten context necessary to understand, if enough vector relationships can be teased out of a sufficiently large corpus of data. It could easily be that our own “understanding” is also derived from the vast n-dimensional matrix that represents all human cultural knowledge.
- wongarsu 3y agoI would argue that "LLMs are merely token predictors" is even a misguided argument. They are token predictors because that's how we designed the only viable output, but that doesn't necessarily limit the amount of intelligence inside the model. You can have dumb token predictors, and you can have AGI that is a token predictor. For a though experiment, you could put a human in a room, and only communicate with them via a terminal. You can write them messages, but the only way they can communicate to you is by giving you a probability distribution for the next token of their answer, then you pick one with a method of your choosing, tell them what you picked, and they choose the next probability distribution. This would severely hamper their ability to communicate effectively, and somebody who only sees the interface and doesn't know the probabilities might conclude it's "only a token predictor". But is it any less of a general intelligence because of that? After all it's still a human, just with a "dumb" input/output protocol that happens to be a token predictor.
- sebzim4500 3y agoYou are just arguing over the definition of words. Who cares? Obviously you can define words like 'learn' to exclude AI training but there is no associated benefit to clarity.
- chpatrick 3y agoDefine learn, know and understand please.
- spicyusername 3y agoWhatever definition you choose, the statement is still likely going to stand. LLMs just predict tokens. It's truly amazing how much value you can get out of that, but we don't need to make more out of it than is necessary. These models were trained on an unfathomable amount of text generated by real humans, we shouldn't be surprised that they sound convincing. But we also shouldn't confuse their ability to sound convincing with legitimate intelligence. If anything, it seems like we're learning that the Turing test isn't a good predictor of actual intelligence. Mostly because real humans are really easily tricked into seeing intelligence where there is none (and paradoxically not seeing intelligence where there is).
- honzabe 3y ago>> LLMs just predict tokens. Out of curiosity, how do you think human intelligence works? I am no expert and I don't claim to know how the human brain works but my layman's mental model of thinking would be basically that: biological neural networks that "do" the thinking are token prediction machines, except "tokens" are not words that appear on the screen but thoughts that appear in consciousness... while the underlying machinery is in both cases a network of units that fire (or not) based on the connections between them. Sure, human experience is a lot more than just intelligence (I have no mental model of how consciousness or qualia might work) but it is surprising to me that so many people keep repeating arguments similar to yours - that LLMs do not really think, they are "just" doing [X] (where X usually describes how I imagine human intelligence works) - if there is more to human intelligence, what do you think it is?
- spicyusername 3y agoif there is more to human intelligence, what do you think it is Thousands of dedicated scientists have spent their entire careers trying to answer this question and we still don't have something remotely approximating an answer. If the answer to "what is consciousness" was as simple as some groups of linear algebra equations, I think someone would have figure that out by now. A gut feeling that one thing is kind of analogous to another is just that. A carpenter and a woodpecker both bang on wood, but knowing anything at all about one doesn't really tell you anything about the other. I suspect laypeople would be less likely to try to make the connection between LLM graphs and physical neurons, if they wouldn't have called those groups of equations "neural networks". They share the name, but not much else.
- wongarsu 3y agoHowever they can exhibit behaviors that are indistinguishable from learning, knowing and understanding. I'm in favor of duck-typing these words. If I would have called it learning on a human, I can call it learning on a complex neural networks who's inner workings we don't fully understand.
- xmodem 3y ago> I'm in favor of duck-typing these words Hard disagree. A subset of behaviours resembling or even "indistinguishable" from understanding should not be enough to meet the definition. Maybe some day we will have AI systems that get there, but LLMs ain't it.
- BoiledCabbage 3y ago> Hard disagree. A subset of behaviors resembling or even "indistinguishable" from understanding should not be enough to meet the definition. Nope, it doesn't work this way. One doesn't get to sit and try to say what something isn't w.r.t a 1st party experience (like understanding or consciousness), without saying how a 3rd party can define what it is. There is no valid argument for "this isn't understanding because my gut says so" without giving a way a person (a 3rd party to an llm) can say what understanding is. We will never have a 1st person "perspective" of a machine LLM. So if you want to claim as a 3rd party what understanding isn't, you need to define, using tool available to a 3rd party observer, what it is beyond "when my gut says so". And if you're defining something by external observation, you are duck typing. Put succinctly, the only practical (non- philosophical) definitions of consciousness, understanding, and intelligence will all be done via duck typing. As no one will ever have a 1st person view of these phenomena for other entities.
- Jensson 3y ago> If I would have called it learning on a human But you wouldn't call that learning as a human. If you include a chess manual in the learning corpus of an LLM, but not any examples of playing chess, then the LLM can give you all the rules about chess, but it can't play chess. That isn't learning, that is just parroting. A human who can recite the rules of chess but can't play chess, would you say that he understands the rules of chess? No, he just know how to output the words and not perform the actions, so there is no understanding there. Same with the LLM, nobody can say it understands these things since you would never said a human understood under similar conditions.
- hackinthebochs 3y agoArguments are a lot more useful than assertions. What reasons do you have for these claims? I argue[1] that LLMs do understand in some cases. What's your response? [1]: https://www.reddit.com/r/naturalism/comments/1236vzf/on_large_language_models_and_understanding/ https://www.reddit.com/r/naturalism/comments/1236vzf/on_larg...
- Jensson 3y agoIf you include a manual in its training set then it still doesn't understand the data in the manual. It can recite the manual since it is trained to recite tokens, but it isn't trained to understand the manual so none of the logic in the manual will be encoded in an LLM. For example, if you train an LLM with a chess manual but no chess games then the LLM can recite all the rules of chess, but it wont be able to play chess. That is what people mean when they say "LLM's are just token predictors, they don't understand anything". So it has nothing to do with spirituality or consciousness or metal vs brains, its just a logical argument based on the limitations of just training the model to predict tokens to match data it has seen. Since it only tried to predict tokens it has seen, and it has never seen a chess game, it doesn't matter if it has seen all the rules, it was never trained to try to follow rules it sees so it can't play chess.
- hackinthebochs 3y ago>It can recite the manual since it is trained to recite tokens, but it isn't trained to understand the manual so none of the logic in the manual will be encoded in an LLM. This is just to repeat the initial assertion with more words. What arguments do you have in favor of this claim? The example you give doesn't track with LLM-based agents. And I'm sure I've come across demonstrations of GPT-4 playing a made up game given the rules of the game, but I can't put my hands on it at the moment.
- xpe 3y ago^ For certain definitions of “learning”. The commenter above certainly understands that there are a diversity of meanings involved with the word “learning”.