3 ms·
The contention that there is no grounding because the training data is linguistic and thus can only reference a world model is disproven in "This sentence has f
by ToValueFunfetti 3mo ago
The contention that there is no grounding because the training data is linguistic and thus can only reference a world model is disproven in "This sentence has five words"- there's real, grounded information about what "five" means within that sentence. While that's a trivial counterexample, I don't know that it's an obvious one (I didn't come up with it myself).
It's not a criticism of the paper itself, but multimodal models came shortly after and provide grounding that is more of the sort the paper is getting at, and it didn't seem like anybody updated on that at all. If multimodal models were still stochastic parrots by the original argument, humans would have to be as well; we don't have any way to ground anything beneath sense data and evolution can't have programmed some innate grounding into us because it didn't either. But (and maybe this is my own misperception) nobody threw in the towel at that point.
I confess I never read the original paper until now, opting to absorb by osmosis instead, and I was quite surprised that they don't really make a deeper case than that. After just a few paragraphs about how they can't be grounded because humans don't express their thoughts directly, it lurches into a page about how they can be biased by training. And they certainly can be, but that has little to say about their stochastic nature- humans are biased as a rule with no exception. (For the record, I only read the Stochastic Parrots section before this reply.)
It's not really a bad paper, but I don't see why it ever carried the esteem it did. Hating on it is like hating on Taylor Swift- she's fine, yes, but for her level of success, one is inclined to question every dumb lyric where others get a pass. (Apologies to Swift fans, substitute a successful artist you don't care for here.)
- dwa3592 3mo ago>>The contention that there is no grounding because the training data is linguistic and thus can only reference a world model is disproven in "This sentence has five words"- there's real, grounded information about what "five" means within that sentence. did you think this through? imagine the sentence was "This sentence has four words", now extrapolate that to all the shit that can exist in a dataset and train a model on that dataset - do you know what will happen? - go ahead and think it through.
- ToValueFunfetti 3mo agoI don't think this tone is at all justified. If you think otherwise, I do ask that you point out where I went too far in a comment that I feared was overburdened by caveats and admissions of my own human flaws. "This sentence has five words" is going to appear far more often than "This sentence has four words". This is the entire premise of LLMs working at all, stochastic parrots or otherwise.
- dwa3592 3mo agoyou are right, i was more curt than i should have been. apologies. but you helped prove my point: >>"This sentence has five words" is going to appear far more often than "This sentence has four words". it's not about this at all. your point is about data quality. you need to take a step back. the point is that if you trained a language model just on this data set which has sentences akin to "this sentence has two words" - the model is going to learn that. this shows that the language modeling itself doesn't truly provide an understanding of the real world. you can train a language model with the most advanced technology on shitty data and the model will start providing shitty outputs confidently - the model will never say "hey there is something wrong with the data i am trained on". thats what 'understanding of the world' meant in that paper.
- ToValueFunfetti 3mo agoIf someone substituted all of your sensory inputs for something else for your entire life, how would you notice? If you wore contacts from birth that made the sky red and earbuds that censored when people said it was blue, on what basis would you realize that was wrong? I don't see what this says about the architecture of your brain, and I don't think it's the point being made in the paper. That the training data must statistically connect to reality in order for the model to model reality doesn't seem that important.
- dwa3592 3mo ago>>If someone substituted all of your sensory inputs for something else for your entire life, how would you notice? I don't know how or if I will notice. That's the biology, chemistry and physics of the brain that I don't know. I hope someone is looking into it. But this does not mean in any way that LLMs are similar to our brains!!!!!! WE DON'T KNOW HOW OUR BRAINS WORK. So going back to language modeling - language models were stochastic in nature when that paper was written, they still are albeit we are trying to make them as deterministic as possible.
- fluoridation 3mo agoTo add to dwa3592's comment, a sentence is not a self-contained idea. The sentence doesn't include what any of the words in it mean, nor what "this sentence" refers to. The exact same sentence can mean different things depending on the text that surrounds it.
- ToValueFunfetti 3mo agoFair point, and on its own it would be surprising to learn what "five" means from that sentence. But you can extrapolate- across a billion sentences, there will be "the next sentence has five words"s and "this sentence are grammared wrong" and so on. It would not be at all impossible to ground a world model on pure text for that reason. And 'not impossible' is sufficient to invalidate the paper's argument.
- fluoridation 3mo agoLet me provide a less superficial response, then. >If multimodal models were still stochastic parrots by the original argument, humans would have to be as well; we don't have any way to ground anything beneath sense data Animals don't passively learn from their perceptions, don't have a separation between training and inference, and don't have a prompt-response execution model. Besides its fundamental biology, the grounding an animal brain has is that when it outputs a motor signal it receives some feedback as to the effects of a signal of that strength. A bird learns to fly because the grounding truth of aerodynamics and gravity consistently respond in a specific way to the flapping of its wings. It doesn't learn by passively replaying thousands of hours of somatosensory recordings of flights. A multimodal model doesn't have the capacity to do much with a prompt. It has no head to turn to look at an image from a slightly different angle to attempt to gleam more information, doesn't have the capacity to interact with the real thing the image represents in any way, and even if it requests another angle and is given it, it lacks the capacity to learn that new information permanently. A multimodal model knows about images of pipes and facts about pipes, but doesn't know pipes; it doesn't have literally first-hand experience with them. >evolution can't have programmed some innate grounding into us because it didn't either. What do you mean? Of course genetics programs ground truths. For example, "forward is the way your face points when your neck is relaxed" and "if you can feel it, then it's part of your body".