6 ms·
Fair point, and on its own it would be surprising to learn what "five" means from that sentence. But you can extrapolate- across a billion sentences, there will
by ToValueFunfetti 3mo ago
Fair point, and on its own it would be surprising to learn what "five" means from that sentence. But you can extrapolate- across a billion sentences, there will be "the next sentence has five words"s and "this sentence are grammared wrong" and so on. It would not be at all impossible to ground a world model on pure text for that reason. And 'not impossible' is sufficient to invalidate the paper's argument.
- fluoridation 3mo agoLet me provide a less superficial response, then. >If multimodal models were still stochastic parrots by the original argument, humans would have to be as well; we don't have any way to ground anything beneath sense data Animals don't passively learn from their perceptions, don't have a separation between training and inference, and don't have a prompt-response execution model. Besides its fundamental biology, the grounding an animal brain has is that when it outputs a motor signal it receives some feedback as to the effects of a signal of that strength. A bird learns to fly because the grounding truth of aerodynamics and gravity consistently respond in a specific way to the flapping of its wings. It doesn't learn by passively replaying thousands of hours of somatosensory recordings of flights. A multimodal model doesn't have the capacity to do much with a prompt. It has no head to turn to look at an image from a slightly different angle to attempt to gleam more information, doesn't have the capacity to interact with the real thing the image represents in any way, and even if it requests another angle and is given it, it lacks the capacity to learn that new information permanently. A multimodal model knows about images of pipes and facts about pipes, but doesn't know pipes; it doesn't have literally first-hand experience with them. >evolution can't have programmed some innate grounding into us because it didn't either. What do you mean? Of course genetics programs ground truths. For example, "forward is the way your face points when your neck is relaxed" and "if you can feel it, then it's part of your body".
- ToValueFunfetti 3mo agoA bird doesn't learn gravity or aerodynamics, it has no 'sense of physics'. It has sensory neural activity that a scientist can show is tied to these things, but you could, at least in theory, falsify the entire experience of the bird. There is nothing in a bird's brain that directly percieves reality. Colors and sounds and textures and so on are all false primitives that don't exist in nature without us, if that's clarifying. And because evolution acts through organisms, its only access to base reality is through their senses. After thousands of years of effort, we have some pretty good models of what reality actually is, but they're still not 'grounded' in the sense you describe.
- fluoridation 3mo ago>A bird doesn't learn gravity or aerodynamics, it has no 'sense of physics'. That's not what I said. What I said was that it's physics that provides the ground truth. >you could, at least in theory, falsify the entire experience of the bird It wouldn't be a bird anymore, but a dysfunctional cyborg with false perceptions. >There is nothing in a bird's brain that directly percieves reality. Yes, of course there is. Animal sensory organs do not produce false information, nor do they provide the brain an interpretation of what they perceive. The brain may fail to distinguish hallucinations from reality, but those are processes internal to the brain. What it gets from the body is raw physical measurements, and what it sends out is raw motor commands. A multimodal model doesn't have the same direct access to reality, it just has collections of words and images. It has no capacity to determine the reality of a photograph of a sunset or a CGI render of a dragon. The word "real" is itself meaningless to it; they're both real in that they appear in its training corpus. It lacks the capacity to investigate these stimuli in any way, and can just learn to associate different stimuli in arbitrary ways that appear to make sense to us, but nothing else.
- ToValueFunfetti 3mo ago>It wouldn't be a bird anymore, but a dysfunctional cyborg with false perceptions. The criticism in the paper is of the architecture of LLMs, isn't it? The paper contends """Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. [...] an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot""" They're saying that the model cannot learn anything about reality irrespective of training data. Your point is an interesting one, but I think it's distinct. To your point though, this is unfalsifiable from the perspective of the "bird". I can't prove that I'm not a dysfunctional cyborg with false perceptions, which makes me wonder if that's a meaningful distinction. >Animal sensory organs do not produce false information, I happen to be in possession of some of these and I think this needs a "usually, under ordinary conditions". (Nitpicky and not critical to my point, but I liked the beginning of this sentence too much to edit it out) >nor do they provide the brain an interpretation of what they perceive Whereas this I'd argue is not true at all. My cones interact with wave-particle photons at particular wavelengths. I can't even conceptualize wave-particle duality (though some humans can), but "red" and "blue" are the bread and butter of my visual consciousness. These correspond to firings of my sensory neurons much more than they correspond to anything in reality. If that's not interpretation, what is it/what is interpretation? >The word "real" is itself meaningless to it; they're both real in that they appear in its training corpus. I expect we agree that I can show a multimodal model a real and a CGI picture and it can tell me which is which. I can take a CGI dragon to GPT 5 and say "look what I found in my backyard" and it will say "Yeah right". Are you saying this is something only possible thanks to RLHF or other modern techniques? That may be the case, unsure how to test that without access to pretraining-only models. Or would you say my experiment is faulty here and doesn't get to your underlying claim? On the flipside, I could show some meh drawings of fairies to Arthur Conan Doyle, and he'd say "Whoa, this changes everything". I consider him to be one of the great rational minds of history, but he was unable to pass your test here. (In fairness, he was in his 60s and his senses may have dulled, though his belief in spiritualism at large dates to his prime). Appreciate the conversation!
- supern0va 3mo ago>and even if it requests another angle and is given it, it lacks the capacity to learn that new information permanently. I'd argue this isn't true today, but that the loop for incorporation is long (ie, the next training or finetuning run). >A multimodal model knows about images of pipes and facts about pipes, but doesn't know pipes; it doesn't have literally first-hand experience with them. Wouldn't this mean that any human who hasn't seen a pipe in person or interacted with it, similarly doesn't "know" a pipe? Most of us haven't interacted with the vast majority of "things" in the world, yet we're still able to build a model and abstractions for them such that we can reason about them, right?