5 ms·
It's funny the author wasted energy composing this after admitting he barely knows the origin of the sentence. Chomsky invokes it in "Syntactic Structures" to i
by 0xddd 6y ago
It's funny the author wasted energy composing this after admitting he barely knows the origin of the sentence. Chomsky invokes it in "Syntactic Structures" to illustrate that the grammaticality of a given sentence doesn't fully explain the odds of it appearing in a large corpus. "Furiously sleep ideas green colorless" is another low probability sentence, yet a native speaker couldn't perform these sorts of mental gymnastics to twist some meaning out of it.
- gizmo686 6y agoI couldn't even get through all of the original article. Having said that; and having formally studied linguistics at an undergrad level, "Furiously sleep ideas green colorless" is different than "Colorless green ideas sleep furiously" in that the former is not grammatical. It is also the case that non-grammatical sentences can have meaning. For example, I believe that most English speaker would agree that the sentence "me happy" is not grammatical, and indicates that the speaker is happy.
- 0xddd 6y agoRight. The sentences are both very unlikely to be uttered even though one is grammatical and the other is not, so there must be another factor in play.
- canjobear 6y agoChomsky was arguing that probability is useless for defining and studying grammaticality. I'm not so sure. GPT-2 says log P("Colorless green thoughts sleep furiously.") = -53.64797019958496 log P("Furiously sleep thoughts green colorless.") = -65.46656107902527 The ungrammatical one is lower probability. But those are famous sentences, and probably present in the training data, so let's try log P("Colorless blue ideas hibernate angrily.") = -60.12953460030258 log P("Angrily hibernate ideas blue colorless.") = -70.02637100033462
- 0xddd 6y agoI think the more interesting result (and more relevant to Chomsky's point) would be to work in the other direction. If you instead produce a list of sentences with similar log probabilities you will see that it contains a mix of grammatical and ungrammatical utterances. This implies something more is needed to distinguish them.
- canjobear 6y ago> If you instead produce a list of sentences with similar log probabilities you will see that it contains a mix of grammatical and ungrammatical utterances. Yes, Chomsky mentions this in a footnote. But as far as I know, it hasn't been tried with modern language models. There's been some interesting work that tries to reproduce grammaticality judgments in terms of language model probability after controlling for length and lexical content. It turns out it works pretty well. For instance https://arxiv.org/pdf/1910.14659.pdf https://arxiv.org/pdf/1910.14659.pdf
- 0xddd 6y agoI wish there were a freely available copy online, I could link, but the passage is at the end of chapter 2 of Syntactic Structures. It's not a footnote, but rather the crux of his argument, I believe: > "... a structural analysis cannot be understood as a schematic summary developed by sharpening the blurred edges in the full statistical picture. If we rank the sequences of a given length in order of statistical approximation to English, we will find both grammatical and ungrammatical sequences scattered throughout the list; there appears to be no particular relation between order of approximation and grammaticalness. Despite the undeniable interest and importance of semantic and statistical studies of language, they appear to have no direct relevance to the problem of determining or characterizing the set of grammatical utterances. I think that we are forced to conclude that grammar is autonomous and independent of meaning, and that probabilistic models give no particular insight into some of the basic problems of syntactic structure." I do think it's an important point for people to recognize. Scientific theories don't arise on their own out of large-scale statistical analyses. There is a lot of faith being put in deep learning methods these days, which are great for prediction, but not inference.