14 ms·
IMO Chomsky (and Marcus) are mostly wrong on LLMs. Quoting the summary of a post[1] I wrote recently: - LLMs are sometimes said to be “just” shallow pattern ma
by erwald 4y ago
IMO Chomsky (and Marcus) are mostly wrong on LLMs. Quoting the summary of a post[1] I wrote recently:
- LLMs are sometimes said to be “just” shallow pattern matchers, “just” massive look-up tables or “just” autocomplete engines. These comparisons amount to a form of (methodological) reductionism. While there’s some truth to them, I think they smuggle in corollaries that are either false or at least not obviously true.
- For example, they seem to imply that what LLMs are doing amounts merely to rote memorisation and/or clever parlour tricks, and that they cannot generalise to out-of-distribution data. In fact, there’s empirical evidence that suggests that LLMs can learn general algorithms and can contain and use representations of the world similar to those we use.
- They also seem to suggest that LLMs merely optimise for success on next-token prediction. It’s true that LLMs are (mostly) trained on next-token prediction, and it’s true that this profoundly shapes their output, but we don’t know whether this is how they actually function. We also don’t know what sorts of advanced capabilities can or cannot arise when you train on next-token prediction.
- So there’s reason to be cautious when thinking about LLMs. In particular, I think, caution should be exercised (1) when making predictions about what LLMs will or will not in future be capable of and (2) when assuming that such-and-such a thing must or cannot possibly happen inside an LLM.
https://www.erichgrunewald.com/posts/against-llm-reductionism/ https://www.erichgrunewald.com/posts/against-llm-reductionis...
- resters 4y agoThink of it this way: They create plausible patterns of language that seem (to humans) similar to human-generated language. That is what they are "good at". Imagine an "AI" that "knows" about rhyme and meter and can generate a seemingly infinite collection of catchy songs that follow the II V I chord progression, only the sequence of words doesn't necessarily make any sense. LLMs are a more powerful version of that kind of thing that does put the words in a sequence that feels like natural human language. Since they are trained on information, the sequences chosen often seem informational, even though they are just patterns. It's really a lot like autocomplete. In other words, a lot of useful information may accidentally be encoded into the model, but so is a lot of stuff that looks/feels similar but is nonsense.
- erwald 4y agoI think I mostly agree with you, but I think this framing is a bit misleading. On autocomplete, I'll just lazily quote the relevant part from my post: It’s completely true that LLMs are trained on next-token prediction (although some, like ChatGPT, are then additionally trained using reinforcement learning with human feedback). It’s also completely true that this fact profoundly influences the texts they generate. So I don’t think it’s unreasonable to call LLMs autocomplete engines or to emphasise next-token prediction. But I think it’s subtly misleading: - Though LLMs were trained to optimise success on next-token prediction, that is not necessarily what they do. We don’t know what it is they do. The training process reinforces behaviours/heuristics in the model that tend to cause it to make better next-token predictions on in-distribution data. This does not mean that those behaviours/heuristics are fundamentally “about” optimising next-token prediction, especially when the model encounters out-of-distribution data. - The usual example here is human evolution. Humans were shaped by a process that optimised for reproductive fitness. This gave us a bundle of drives such as family kinship, prestige and sexual pleasure – drives that aren’t fundamentally about optimising for reproduction, which becomes evident as we enter a new environment – one with contraceptives, say. - Optimising for a task for which intelligence is useful encourages the optimised thing to become more intelligent. Sam Altman gave expression to this last week when he wrote, “Language models just being programmed to try to predict the next word is true, but it’s not the dunk some people think it is. Animals, including us, are just programmed to try to survive and reproduce, and yet amazingly complex and beautiful stuff comes from it.” - The forms of intelligence that are useful in doing next-token prediction are different from those that are useful in human reproduction, but I think there’s a considerable overlap, as (1) some fundamental abilities, for example using and applying concepts, just seem very broadly useful and (2) the data LLMs are trained on are written by humans, for humans and often about humans and things that matter to us. I think the "they're just autocomplete" take also hides other properties of LLMs, like them seeming to (as mentioned in another comment, and in the post) contain and use world models, and being able to learn general algorithms.
- naasking 4y ago> Since they are trained on information, the sequences chosen often seem informational, even though they are just patterns. The point is is that dismissing it because it's "just patterns" ignores the fact that nobody has proven that human cognition isn't also "just patterns". If it is, then that undermines the whole justification for the dismissal. So the dismissal assumes the conclusion until that evidence is presented.
- orwin 4y agoI think they're more than that, but I also think they're bullshit generators. Informed bullshit generators even. Which is quite useful to do some tasks: writing fiction (it now wrote at least 6 characters in multiple rpg campaigns I participate in) and probably others (I used it to write tests data, but in fine I had to revalidate it by hand. Could do better.) I do not see LLMs be more than that in the near future though. More informed, more accurate bullshit generators? Yes probably. And I'm not saying it isn't impressive and very close to human behavior. I still think we're missing a part in our models though, and until we find it, breakthrough will be limited to improvements on accuracy and information.
- erwald 4y agoThey definitely hallucinate a lot too. But they also seem to do things genuinely like reasoning, e.g. they seem to contain and use "cognitive" world models, and seem to be capable of learning fully general algorithms. Beyond that, many intellectual capabilities are also useful for successful bullshitting, so even if we can confidently say that that's all they do (in some sense), that doesn't mean they don't also do something-like-human-reasoning etc.
- orwin 4y agoThat's basically what i was saying. It's close to human capability. I don't really know how to explain it, but to me we have at least two kind of reasoning, swallow and deep. LLMs could well be capable of the first.
- JohnAaronNelson 4y agoCan you explain to me how choosing “the most likely next word” is reasoning? Can there be reasoning without language?