3 ms·
> I think the phrase "stochastic parrot" is misleading and has fooled millions of people into thinking that LLMs can't do genuine reasoning about situations the
by bobsomers 2y ago
> I think the phrase "stochastic parrot" is misleading and has fooled millions of people into thinking that LLMs can't do genuine reasoning about situations they've never seen before, nor been trained on, which is wrong because LLMs definitely are doing genuine intelligent reasoning.
Can you describe specifically which part of an LLM architecture does the "reasoning" and how it works? Because every architecture I'm familiar with is literally just a fancy way of predicting the likelihood of the next token given the stream of previous tokens and the known distribution of token based on the training data. This is not reasoning. This is simple statistical prediction, making the "stochastic parrot" analogy actually quite accurate.
> What model training is doing is building up a semantic space of vectors from which astronomically large numbers of true facts and ideas can be derived during inference. I mean like a number of facts larger than the number of molecules in the known universe. A googolplex more facts than the sum of all of humanity has ever "thought".
Sort of. It's like an extremely lossy compression process which has absolutely no guarantee that "truth" was maintained in the process. Also I'm extremely dubious of your claim that it can accurately encode "a number of facts larger than the molecules in the known universe" given that it's trivially easy to get the best LLMs to give you an incorrect answer to a question any human would easily get right.
- quantadev 2y agoThe "reasoning" is an emergent property that no one understands yet. Yes we do all the training only to try to predict the next word (i.e. train to do word prediction), yet with enough training data, then at some scale (GPT 3.5ish) the embedding vectors in semantic space begin to build a geometric scaffolding in into the weights, for lack of a better way to phrase it. If you know about facts like (Vector(man) minus Vector(woman) equals Vector(king) minus Vector(queen)), that's an indication that this "scaffolding" is taking shape. It means the "concept of gender" has a "direction" in the roughly 4,000 dimensional vector "space". This vector behaves geometrically, so that vectors behave in vector space as if it was a geometric space of sorts (there are directions and distances), even though there's no true space coordinates, just logical "directions". Mankind doesn't quite yet understand the "Geometry of Logic". LLMs prove we don't. I think it's a new math field to be invented. As far as the actual number of "facts" contained in an LLM, I think you have to consider something that's a function of the number of bits in an entire model, and ask how many "states" can that store, as a rough approximation from an entropy standpoint. But these aren't pure facts. They're reasoning. I guess you can call reasoning something like "fuzzy facts", so there's a bit of uncertainty to each one of them. LLMs don't store facts, they store fuzzy reasoning. But I call it "factual" when an LLM fixes a bug in my code, or correctly states some piece of knowledge.
- bobsomers 2y agoPerhaps we disagree on semantics here, but IMHO I wouldn't call this "reasoning". It's essentially just data compression, which is exactly what you get by constructing an encoder network that minimizes loss while trying to maximally crunch down that data into a handful of geometric dimensions. "Mankind doesn't quite yet understand the geometry of logic" is laying it on a bit thick with the marketing speak, IMHO. It's just data compression whose result is somewhat obvious given what the loss function is optimizing for. If a structure capable of real reasoning was being built, I wouldn't expect LLMs to get tripped up by simple questions like "How many Rs does the word Strawberry have in it?". There are only two simple reasoning systems you need to solve this question. You need to learn the English alphabet and you need to be able to count to a handful of single digit numbers, both tasks that kids of age 3-4 have mastered just fine. Putting together these two concepts allows you now to reason your way through any such question, with any word and any letter. Instead, LLMs perform how we would mostly expect a stochastic parrot to react. They hallucinate an answer, immediately apologize when it's called out to be wrong, hallucinate a new, still incorrect answer, immediately apologize again, until they eventually get stuck in a loop of cursed context and model collapse. I'm not suggested that an LLM couldn't learn such a reasoning task, for example, but it would need to look at many training examples of such problems, and more importantly, have an architecture and loss function that optimized for learning a mechanical pattern or equation for solving that kind of problem. And in that regard, we're very, very far away from LLMs that can do any kind of generic reasoning, because I haven't seen any evidence that those models are generic enough that you can avoid learning lots and lots and lots of specific ways to approach and solve problems. One thing I think it's critical to keep in mind is that improvisation upon contextually relevant data in your compressed knowledge base is not reasoning. It might sound convincing to a human reader, but when it's failing at much simpler reasoning tasks the illusion really is shattered.
- quantadev 2y agoWhen people say "If it was reasoning, then it would be able to know, How many Rs does the word Strawberry have in it?", but that's not quite right, but I would say this instead "If it was reasoning THE SAME WAY HUMANS reason....then it would be able to...". Humans do reasoning a certain way. LLMs do reasoning a different way. But both are doing it. But since it's not reasoning the way people do (but very differently), yes it can make mistakes that look silly to us, but still be higher IQ than any human. Intelligence is a spectrum and has different "types". You can fail at one thing but be highly intelligent at something else. Think of Savantism. Savants are definitely "reasoning" but many of savants are essentially mentally disabled by many standards of measurement, up to and including not being able to count letters in words. So saying you don't think LLMs can reason, and giving examples fails as evidence of that, is just a kind of category error, to put it politely. The fact that LLMs can fix bugs in pretty much any code base shows it's definitely not doing just simple "word completion" (despite that way of training), but is indeed doing some kind of reasoning FAR FAR beyond what humans can yet understand. I have a feeling only coders truly understand the power of LLMs reasoning because the kind of prompts we do absolutely require extremely advanced reasoning and are definitley NOT answerable because some example somewhere already had my exact scenario (or even a remotely similar one) that the model weights essentially had just 'compressed'. Sure there is a compression aspect to what LLMs do, but that's totally orthogonal to the reasoning aspect.