19 ms·
“Just a statistical text predictor”
- CmdrLoskene 3y ago"If it looks like magic to me, it must be magic."
- readyplayeremma 3y agoI understand that you may believe your ability to think is magic, but I assure you, you are also just a statistical predictor, and not magic. Thinking and cognition themselves, beyond the medium they happen within, are moving out of the realm of philosophy and into the realm of science now. It just so happens that quite a lot of people are uncomfortable with the idea of not being as special and unique in this way anymore.
- djtriptych 3y agoYes. I think this is exactly right and feel like I'm waiting for people to get over the shock. It's really interesting (but yes, initially scary) to think about what really makes us human. In a sense we _get_ to think about all the things we can do that robots can't. I hope at some point most feel a sense of relief that a robot can assist in the robotic things we're often forced to do.
- tgv 3y agoYou're positing this with way too much confidence, like GPT. We absolutely do not understand what thinking is. As far as I'm concerned, we are at least self-correcting, conscious, aware statistical predictors. It's not only prediction either. Some of it is direct response or autonomous behavior (your heart doesn't beat because it tries to predict something).
- dS0rrow 3y ago> I understand that you may believe your ability to think is magic. Nobody said that. I am sick of the false dichotomy of being given only two choices: a) It's magic b) It's akin to our latest discovery
- pupperino 3y agoA lot of people are quite unhappy with it, yes, but... Consciousness is still very much a philosophical problem. Think of Marr's levels, functionally, LLMs and humans are quite similar, but algorithmically? How about the implementation layer? There are still many, many mysteries to be solved in the realm of philosophy of mind and LLMs should serve as reminders that Consciousness and Intelligence are not the same. May I suggest this reading? https://qualiacomputing.com/2022/06/19/digital-computers-will-remain-unconscious-until-they-recruit-physical-fields-for-holistic-computing-using-well-defined-topological-boundaries/ https://qualiacomputing.com/2022/06/19/digital-computers-wil...
- datathrow0007 3y agoNot a fan, OP. No, reducing LLMs to just “statistical text predictor” is an absurdity; but so is leisurely stringing together observations from unstructured experimentation to form a “matter-of-fact” conclusion. Read the ChatGPT paper and specifically hone in on “transformer” and “input processing” (paraphrased). It will give you a clearer picture of what it actually is, rather than appears to be.
- readyplayeremma 3y agoThis is a good overview of a number of emergent things that remain unexplained about the capabilities of some large language models: https://www.scientificamerican.com/article/how-ai-knows-things-no-one-told-it/ https://www.scientificamerican.com/article/how-ai-knows-thin... You can follow the multitude of things from there to the papers published that describe each of the cases in greater depth.
- sebzim4500 3y agoWe do not expect just knowing the differential equations governing biological neurons to give us much insight into human intelligence. Why do you think it should be different for artificial ones? I've read the relevant papers, I've even implemented some of them, and I do not think any of this gave me a better understanding of it's capabilities than playing with it as a black box for a few weeks.
- FPGAhacker 3y ago> We do not expect just knowing the differential equations governing biological neurons to give us much insight into human intelligence. No, I expect not. But I think it’s a mistake in thinking and loose use of language to say that differential equations, or any mathematical representation, governs anything. They are models of whatever may actually be happening, and are necessarily modeling a simplification of the behavior observed. Don’t mistake the model for reality.
- sebzim4500 3y agoFair enough, but I would argue that even having perfect knowledge of everything happening in a biological neuron down to the quantum level would not help very much when you are trying to understand how the brain makes decisions.
- Spivak 3y agoThis article could really could be shortened as Calling Llms a statistical text predictor is like calling a brain a bunch of neurons responding to electrical impulses. It's true but misses the thing we care about which is the emergent behavior of such systems, in humans we call it consciousness. For Llms which are akin to a language processing center all by itself conciseness isn't the word but a yet unnamed thing.
- jwlake 3y agoAny sufficiently advanced statistical predictor is magic?
- skybrian 3y ago> it is very rare for GPT4 not to be able to understand why it went wrong. That's not what's going on. GPT4 can't know the real reason why it wrote anything. It can never tell you the real reason why it made a mistake. Instead, we can say that GPT4 is fairly good at inventing plausible explanations about why someone did something, but it's no more accurate explaining its own mistakes than it would be at explaining one of your mistakes. [1] A better explanation of what happened is that asking for an explanation resulted in chain-of-thought reasoning, and by telling it that its previous answer was wrong, it was biased to pick a different answer. But this is often fragile. You'd want to do the same test multiple times, and compare regular chain-of-thought reasoning with what happens when you explicitly tell it that a particular answer is wrong. [1] https://skybrian.substack.com/p/ai-chatbots-dont-know-why-they-did https://skybrian.substack.com/p/ai-chatbots-dont-know-why-th...
- famouswaffles 3y ago>That's not what's going on. GPT4 can't know the real reason why it wrote anything. Boy do I have news for you. People can't recreate previous mental states, it's all plausible post-hoc rationalization. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/ And the rationalizations aren't always grounded either. The brain is just fine making up completely bogus explanations you believe to be true but couldn't possibly be so. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522/ https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf The conscious self is at best a co-pilot. A Co-pilot that has shockingly little insight into the operations of the plane but thinks he knows much more. Even your senses gets some pretty big post-processing by your brain you're consciously unaware of. For example, ever wondered why the first second seems longer when you suddenly look at a clock ? Human beings may not have the best visual acuity, but we are quite good at distinguishing colors and exceptional at processing visual information. A significant portion of our brain is dedicated to processing the visual stimuli from our binocular vision. Nevertheless, binocular vision necessitates that both eyes focus on a single point. Our eyes can track a focused target as it moves, but they are unable to smoothly scan across our visual field. This is too difficult to process, so our eyes naturally jump from one point to another. Yet, during these jumps, known as "saccades," our brain does not even attempt to comprehend the blurry sensations that come in. In effect, you temporarily go completely blind! However, you don't feel like you went blind, right? Your vision doesn't go black or anything, and it seems like you could see the entire time. This is because our brains fill in the lag while our eyes are moving, creating the illusion that we were seeing something instead of being blind during that time. Strangely enough, our brains fill it in with *what we see when they stop*! In other words, our brains go back and edit our memory of a split second ago to make it seem like we were looking at our new point of focus the whole time our eyes were moving to actually see it. This phenomenon occurs when you look at the second hand of a clock. When you glance over at the clock, your brain tells you that you were looking at it for longer than you actually were, adding the time your eyes were moving but not yet seeing the clock. If the second hand was moving during this time, your brain will believe the second hand was in its new position longer than it actually was. Based on this flawed memory, it will appear as if the second hand stayed in that position for longer than a second.
- deleted 3y ago[deleted]
- keskival 3y agoLooking at Transformer model architecture and their training schemes to understand how they think misses the point. You wouldn't gain insight into how Doom game works by just looking at schematics of an x86 processor. These systems are Turing complete and they can execute any computation. To see what they actually execute requires looking at what happens in memory while Doom runs, in the activations when the Transformer model executes. At this time we don't have good tools to extract the actual algorithms and submodels these LLMs execute, but there are reasons to believe there are better algorithms learned in there for example for reinforcement learning than what we have ever been able to engineer. See for example my research: https://github.com/keskival/king-algorithm-manifesto#readme https://github.com/keskival/king-algorithm-manifesto#readme
- skybrian 3y agoMechanistic interpretability is just getting started, but there's been some interesting research on induction heads. Here's an explanation: https://www.lesswrong.com/posts/TvrfY4c9eaGLeyDkE/induction-heads-illustrated https://www.lesswrong.com/posts/TvrfY4c9eaGLeyDkE/induction-...
- rcme 3y ago> You wouldn't gain insight into how Doom game works by just looking at schematics of an x86 processor. How is this a relevant analogy?
- arc619 3y agoThat was fascinating, thanks. Please do consider posting this here as its own article, it would be interesting to read others comments. My reading of this - and please correct me if I'm wrong, I'm still learning - is that you're extracting hyper-parameter planes from the data flow in the model's embedding space? Its really exciting to think of the hidden knowledge and relationships we could extract from our own linguistic interactions.
- hamhamed 3y agoIt is a statistical text predictor, and finetuning (RLHF) is the "magic" that gives it meaning and reasoning via weigths
- chrismorgan 3y ago> … though it initially became cautious: > [GPT4]: “equine.” > [Human]; What sort of equine? > [GPT4]: “horse.” The sentence it had been asked to complete ended with “… the resulting animal will look like an”. I suspect “equine” wasn’t the product of caution, but of matching the article an. (There may still be some English accents that use “an horse”—perhaps some Indian and some parts of England—but the vast majority now use “a horse”.)
- bccdee 3y agoI think this misunderstands the reason why people say LLMs are "just" statistical text predictors and misunderstands LLMs in general. Here's where you're wrong: > The real reason behind this bizarre answer is that it matches the form of a more elaborate puzzle where a similar indirect approach is necessary. GPT4 gets caught up in the tendency to match previous examples, and this is a genuine reflection of its cognitive history; it started out as a statistical text predictor. This is not the same as not understanding the question; this is a reflection of a faulty cognitive architecture that can be fixed relatively easily. This is not anything that can be easily fixed. This type of faulty problem-solving is a fundamental property of how LLMs work. It can be attenuated or avoided, but not repaired or removed. LLMs work by assembling language into plausibly-human-written sentences, and if you ask them to engage in problem-solving, they will often write their way to a correct solution, but examples like the 12L/6L jug question are revealing: The LLM doesn't actually reason. It imitates the appearance of reasoning, and does so extremely well, but that's not the same. It doesn't understand why you would pour one jug into another; it's been taught that pouring jugs into each other and arriving at the required number of litres are characteristics of the solutions of similar problems it's been trained on, and so it does all of these things. LLMs imitate human writing, and can, by proxy, reproduce human reason, but they have a tenuous and indirect grasp on it, and this drawback is built into their design.
- oezi 3y agoI see absolutely no reason (except cost) that LLM interfaces shouldn't utilize an internal dialog which is hidden from the user and which can be queried from the first layer of interaction and potentially vice versa. This would emulate much of what we consider human reasoning. And if it quacks like a duck it might be a duck...
- bccdee 3y agoAsking the LLM to be more verbose can bring contradictions to the surface and lead to more reliable answers, but it still doesn't fix the underlying problem. Even when you ask it to show its work, it can still contradict itself, blindly imitate patterns in its training data, and contrive plausible-sounding but false answers. I've got a (rather long) transcript I could post of a simple logic problem which I asked ChatGPT to solve step by step. It started off okay, but then it started completely breaking the rules of the puzzle and the internal consistency of its answer. When I asked it to identify its mistake, it made up a totally different mistake it never made. The thing is, it's not emulating. It's imitating, and there's a big difference: When you emulate a logical process, you're copying it from the bottom up, and the internal logic is the same. If there are mistakes, they're fairly marginal. When you imitate a process, however, you're trying to reproduce the same output without recreating the internal logic that creates that output, and it's easy to detect and extrapolate patterns in the output in ways that don't make sense (like with the 12L/6L jug puzzle).
- mike_d 3y ago> "I think this assessment is wrong on many levels" Ok so in your expert opinion, what is it actually doing? > "Exactly what it does, no one really knows, not even its creators." I mean, it obviously couldn't just be feeding all the previous text into a model that statistically predicts the next likely words and then does some post processing to intelligently pick one of the high ranked words, just like the developers say. Because that would just be a statistical text predictor. People seem to really hate the idea that something so simple can fool them into thinking it is alive, like it is a challenge to their own intelligence.
- sebzim4500 3y ago> I mean, it obviously couldn't just be feeding all the previous text into a model that statistically predicts the next likely words and then does some post processing to intelligently pick one of the high ranked words, just like the developers say. Of course that's what it's doing, but if I ask you how you typed that comment you wouldn't say "trillions of synapses whose connection strength was determined based partly on my genetics but mostly on all my sensory data since I've been alive ended up firing in order to make my fingers press the 'reply' button". Do not mistake understanding the substrate and the loss function with understanding the actual model.
- mike_d 3y agoSure, but saying "it is magic and nobody really knows how" when we have a solid scientific basis for how messages are carried from the brain to the fingers is highly disingenuous.
- sebzim4500 3y agoI agree that saying that "it is magic and nobody knows how" is not a good explanation for anything. But saying "we understand the substrate so we are done" is no better.
- mikojan 3y ago> One of my first engagements with GPT4 was to discuss the concept of maralia. [...] I chose maralia because the discussion involved a complex philosophical concept that I knew was not in its training set. [...] GPT4 not only understood the idea, but it was able to use the concept appropriately [...] This is like the time when Alan Sokal submitted an article to Social Text but this time it is the social scientist playing a trick on themself.