16 ms·
Modern language models refute Chomsky’s approach to language
- hackandthink 3y agoIf you got time for 48 pages.
- hackandthink 3y agoActually only 30 pages text. Many, many references.
- mistertiddly 3y agoNo, they don’t.
- HPsquared 3y agoEmpirical vs analytical
- renewiltord 3y agoI only skimmed a few pages, but it doesn't appear to be about the Chomsky hierarchy - which was what I was curious about.
- cma 3y agoDeep Mind had this paper on that, showing weaknesses of transformers relative to other models in generalizing from example productions from different tiers of the hierarchy: https://arxiv.org/abs/2207.02098 https://arxiv.org/abs/2207.02098
- nhatcher 3y agoChomsky did quite a lot of things in linguistics. This is about whether or not there is a universal grammar: https://en.wikipedia.org/wiki/Universal_grammar https://en.wikipedia.org/wiki/Universal_grammar There is a fairly interesting read from Steven Pinker: "The language instinct" about this topic.
- mcguire 3y agoWhile I think some of the points in the article are interesting, the usual evidence for the Chomskian approach is the relative lack of input data for learning language by children in the wild. How much input data is used to train modern language models?
- api 3y agoDo children really lack input data though? Human sensory input is quite a lot of data. Our languages may have common structure because that structure reflects causality and physics encountered by directly sampling the real world.
- sudosysgen 3y agoAnd how exactly is the structure of the world transformed into a structure of language? Chomsky would say through a universal grammar, which is a framework through which you can set up ways to encode the structure in the world as words.
- api 3y agoThis implies that the combinatorial space of possible grammars is large. If it isn't, then the structure of the universe would imply a narrow set of possible workable grammars. The combinatorial space of languages is obviously infinite, but is this so for grammars? If not then you would expect many languages to share the same grammars. Seems analogous to the same question about mathematics. Are there many different possible arithmetics? No. There are infinitely many ways to express arithmetic symbolically but there is only one arithmetic. 2 + 2 never equals 5.
- deleted 3y ago[deleted]
- cma 3y agoYou have Hellen Keller who went blind and deaf at 19 months.
- neosat 3y agoThe author (bafflingly) seems to have completely missed the point- since anything they state up to page 15 (at which point I stopped reading) does not refute Chomsky's points at all. The author talks about LLMs and how they generate text and then goes on to talk about how it refutes Chomsky's claim about syntax and semantics. However it does not since Chomsky's primary claim is about how HUMANS acquire language. The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. [edited page = 15]
- CydeWeys 3y ago> The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. Also, the LLMs are cheating! They learned from us. It's entirely possible that you do need syntax/semantics/sapience to create the original corpus, but not to duplicate it. Let's see an AlphaZero-style version of an LLM, that learns language from scratch and creates a semantically meaningful corpus of work all on its own. It's entirely possible that Chomsky's mechanisms are necessary to do so.
- famouswaffles 3y ago>Also, the LLMs are cheating! No...they aren't. Humans aren't learning from thin air by any stretch of the imagination.
- deleted 3y ago[deleted]
- adamnemecek 3y agoI'm betting on learning (both human and ML) being analogous to renormalization in QFT. Essentially, you are trying to explain a process in terms of imperfect patterns. With incoming data, you update both your patterns and your locations where these patterns occur at once. There is some research into this https://phys.org/news/2022-05-renormalization-group-methods-brain.html https://phys.org/news/2022-05-renormalization-group-methods-...
- mchaver 3y agoI am agreement with the author that Chomsky is overly negative about LLMs, but I don't think this is a strong refutation (admittedly I've only spent 20 minutes skimming the article). LLMs give us a lot of insight into how computers/algorithms can manipulate and understand language, but that still does not tell me how humans do it (I am open to the possibility that this is how humans do it, but that evidence is not presented here). One of Chomsky's main arguments is the poverty of stimulus (children seem to learn language with relatively little input). Here is what the author has to say: > Large language models essentially lay this issue [poverty of stimulus] to rest because they come with none of the constraints. Modern language models refute Chomsky’s approach to language that others have insisted are necessary, yet they capture almost all key phenomena. It will be important to see, however, how well they can do on human-sized datasets, but their ability to generalize to sentences out-side of their training set is auspicious for empiricism. That doesn't look like a refutation to me yet. We still need to do that test, but that still just tells us how you can do it algorithmically.
- Guidii 3y agoI'm unconvinced. The author seems to be stating that because llms have generated "interesting" statements that all of Chomsky's approaches are invalidated. This seems to assume that the llm isn't somehow incorporating Chomsky's generative grammar system into the model, which would be a very hard thing to prove. Also, the author's stated proof contains flaws. Take for example page 16: Generate ten other sentences like "colorless green ideas sleep furiously": Purple fluffy clouds dream wildly. Blue glittery unicorns jump excitedly. The author claims that the llm has completed this task successfully, completely ignoring the example's pairwise incompatible terms (colorless green, and sleep furiously), and accepting that "purple fluffy" is equivalently meaningless. It is not, and the model has clearly failed.
- aap_ 3y agoIt sure is an interesting time for linguistics now. For the first time there now is a second entity that can understand and produce language. Will be exciting to see how the field will transform.
- zvmaz 3y ago> For the first time there now is a second entity that can understand and produce language But does it really understand language? This reminds me of the Chinese room argument [1]. [1] https://en.wikipedia.org/wiki/Chinese_room https://en.wikipedia.org/wiki/Chinese_room
- aap_ 3y agoOf course this getting into the whole philosophy of AI and consciousness and whatnot. But for practical purposes I would say: yes
- zvmaz 3y agoFrom the conclusion: > First, the fact that language models can be trained on large amounts of text data and can generate human-like language without any explicit instruction on gram- mar or syntax suggests that language may not be as biologically determined as Chomsky has claimed. Instead, it suggests that language may be learned and developed through exposure to language and interactions with others. I'm not a linguist nor a cognitive scientist, but this seems so problematic that I am not sure that I read it correctly. For example, how is the fact that language models "work" contradict the innateness of language in humans?
- taeric 3y agoIt contradicts the idea that you have to teach language using syntax and grammar. Which... I confess I thought was already not believed? We certainly aren't teaching kids in the home how to decline and conjugate words. Are we? (Similarly, languages that have gender are typically just picked up by usage, not necessarily ingrained by reasoning. Which leads to the obvious bad results when people think that there was solid reasoning on those choices, in the first place.)
- pas 3y agoyep. Chomsky used the term "universal grammar", but it's much more likely a universal abstract semantics++ thing (coupled with sound processing stuff, plus repetition, plus a bunch of other stuff in the brain that helps learning in general). how much is innate, what exactly does that mean, all good questions. and of course raw intelligence (pattern matching, strategizing, learning, adaptiveness, modeling, ability to form a sort of consistent and goal-orientedly useful predictive model of the world based on inputs, and goal-oriented control of behavior based on these aforementioned models) by definition can learn language. and of course it's a strange question of are LLMs intelligent in this sense despite lacking goals?
- bbor 3y agoYeah perhaps that’s causing confusion for science enthusiasts new to this debate - Chomsky is definitely not talking about what a lay person would call “Grammar”. Your point is one of his main (implicit) supports though: on the face of it, it seems absolutely insane that a child could pick up complex linguistic concepts in a few short years, with many orders of magnitude less data than a LLM needs to reach the same capabilities
- agnosticmantis 3y agoAlso interesting is Peter Norvig’s essay on the two cultures of language modeling from more than a decade ago: https://norvig.com/chomsky.html https://norvig.com/chomsky.html
- comboy 3y agoTangentially related, but it's interesting that Chomsky stated in a few interviews that he seems LLMs as just plagiarism machines, that they don't create anything new. Which I disagree with - us being creative is also just colliding patterns together. But at the same time I kind of assign higher value to his opinion than mine..
- z3c0 3y agoThese are two different modes being conflated. A person committing plagiarism is akin to a how a GPT creates a document. "Okay, this word... and then this word... and then this word..." This is opposed to modeling a concept in your mind, and then applying language through denotation. This isn't unlike composing a request to be sent over a specific protocol. The data exists independently of the protocol and could even be fit to more protocols, with the right understanding of how to implement them. Sure, you have to read some docs, and maybe use a library somebody else wrote, but nobody in their right mind would call that plagiarism. This is more akin to how language works in the human brain, where each new language is like a different protocol.
- gizmo686 3y agoThis paper completly misses the point of linguistics as a discipline that generative grammar operates in. We have known that it is possible to understand language without innateness. That is what linguists do. If you look at how linguists know about innate features, the answer is almost always by first discovering them explicitly while analysing language data; not by opening a brain to see what is innately inside. [0] The point about innateness is that it takes generations of linguists to learn from a blank slate properties of language that children learn in just years. There are also numerous other arguments for innateness. From the way humans seem to spontaneously develop language in a language deprives environment, to the way language aquasition works being more consistent with other innate behaviours, to the pressence of weird properties that seem to be present across languages for no apparent logical reason. The only insight I see from LLM is the same insight we have seen throughout macine learning. It is not nessasary to understand something if you can throw enough compute at it. This is powerful, and it enables us to do a lot, but it should not be confused with understanding. [0] There are some instances leveraging MRI and other cognative research teqniques to get some insight into the inner workings of human language processing, but their role in developing current linguistics theory is thus far limited.
- zvmaz 3y agoAlso, dogs and cats, and even our close relatives, primates, don't develop a capacity to language.
- azakai 3y agoYes, and that's one of my favorite Chomsky points. I don't remember the exact phrasing, but something like: "Language is innate in humans because in every household practically all children learn it while none of the pets do."
- littlestymaar 3y agoThey[1] do learn some of it, but a fraction so small it doesn't change the argument. [1]: at least the mamal pets, not goldfishes
- ineedasername 3y agoChomsky was describing how humans acquire language. The distinction seems important.
- dr_dshiv 3y agoOne of Chomsky’s main modern claims is that MERGE is the central linguistic operation: the ability to take two elements and merge them into one, hierarchically. https://en.wikipedia.org/wiki/Merge_(linguistics) https://en.wikipedia.org/wiki/Merge_(linguistics)
- michaelhartm 3y agoNobody knows how the LLMs work under the hood. It's just lots of stacked transformers that encode various concepts. Nothing in this book refutes whether Chomsky's concepts are actually being encoded in LLMs or not. For all we know, Chomsky's concept of "binding principles", "binary branching" etc could be represented inside the inner layers of these many billion parameter models. In fact, I'd argue that this is the right research to do. Prove that no transform or feed-forward layer inside the neural net encodes, say "binding principles".
- michaelhartm 3y agoBtw, semantics and syntax is separated in the LLMs (the author is wrong). The embedding function (matmul) can map syntax and the proximity in the embedding (e.g. cosine similarity) is the semantics (that's attention). So not convinced. Chomsky might be wrong or right, but this author hasn't proven it.
- dist-epoch 3y agoThere is a selection effect. Of all possible neural network architectures, so far one of them, Transformers, delivered good results. It's possible this architecture is more similar to the brain language architecture.
- abujazar 3y agoThis is peak bullshit level.
- why-el 3y ago> First, the fact that language models can be trained on large amounts of text data and can generate human-like language without any explicit instruction on grammar or syntax suggests that language may not be as biologically determined as Chomsky has claimed "The fact that this advanced drone released in the year 2080 that can drive with the same agility as an eagle proves that the eagle's flying ability is not as biologically determined as some people claim. In fact, any organism can fly if it sees enough data about flying!"
- pixelmonkey 3y agoOn Chomsky, here is a recording of some audio from Chomsky where he uses a rhetorical argument for why language may be innate to humans rather than (fully) learned. From 1992. Just 5 minutes. https://youtu.be/CPgDALpS-7k https://youtu.be/CPgDALpS-7k Let's make sure what the computer scientists understand Chomsky to have stated is actually aligned. Chomsky didn't say the ONLY way to create language is via the brain. His view, instead, is that evolution programmed language development into the brain -- that it is not learned (entirely) by peer osmosis. That the brain has some structure for language, built-in, which is the unlocked in various ways via socialization. Summary of Chomsky's view, paraphrased: "It is a strange intuition [that most other people have]. Above the neck, we insist everything [in human development] comes from experience. Below the neck, we're willing to accept the idea that [...] it comes from inside. [...] But: it is hard to look at the Sun setting and say, it's not 'setting', the Earth is actually turning. Similarly, with people, it's hard for us to look at them and not see them as minds inside bodies. This leads us to this false approach: below the neck, we are willing to pursue the sciences, and if that leads us to believe development is internally programmed, we'll accept it. But above the neck, we'll be completely irrational; we're going to insist on beliefs and explanations we'd never normally dream of in other rational areas." --- This YouTube clip came to mind, but here is a more detailed explainer from Stanford Encyclopedia of Philosophy: "Clearly, there is something very special about the brains of human beings that enables them to master a natural language — a feat usually more or less completed by age 8 or so. ... This article introduces the idea, most closely associated with the work of the MIT linguist Noam Chomsky, that what is special about human brains is that they contain a specialized ‘language organ,’ an innate mental ‘module’ or ‘faculty,’ that is dedicated to the task of mastering a language. On Chomsky's view, the language faculty contains innate knowledge of various linguistic rules, constraints and principles; this innate knowledge constitutes the ‘initial state’ of the language faculty. In interaction with one's experiences of language during childhood — that is, with one's exposure to what Chomsky calls the ‘primary linguistic data’ or ‘pld’ — it gives rise to a new body of linguistic knowledge, namely, knowledge of a specific language (like Chinese or English). This ‘attained’ or ‘final’ state of the language faculty constitutes one's ‘linguistic competence’ and includes knowledge of the grammar of one's language. This knowledge, according to Chomsky, is essential to our ability to speak and understand a language (although, of course, it is not sufficient for this ability: much additional knowledge is brought to bear in ‘linguistic performance,’ that is, actual language use)." source: https://plato.stanford.edu/entries/innateness-language/ https://plato.stanford.edu/entries/innateness-language/
- JoshTko 3y agoLanguage is essentially compression. And studies show humans automatically compress data relative to chimps. So there are likely biological structures that help humans compress concepts. Compression could also be an emergent property of having more layers in a NN.
- akasakahakada 3y agoJust look how many are still saving the sinking boat named "human-is-special-ism". "Only human can learn and understand language, not other creature in the university, not machine"
- PeterCorless 3y agoThe LLMs don't understand anything. At all. They have rules they respond to. But that doesn't mean they actually "get" a single word of what they're blathering. A computer has no sense of "intent" or "meaning" to what they have indexed and scanned. There is no intent or meaning to what they spit out in generated text. We're back in the Chinese room. https://en.wikipedia.org/wiki/Chinese_room https://en.wikipedia.org/wiki/Chinese_room
- azakai 3y agoThe Chinese Room Argument is considered by many to be incorrect on those matters. But it has led to decades of interesting philosophical debate, which is worth reading to see the various perspectives. This is an excellent summary: https://plato.stanford.edu/entries/chinese-room/ https://plato.stanford.edu/entries/chinese-room/
- anthonybsd 3y agoWhy is everything these days revolve around ChatGPT(etc). You don't need LLMs to refute Chomsky language models. Modern linguistics pretty much rejected [1] his theories on the basis of evidence. [1] https://www.scientificamerican.com/article/evidence-rebuts-chomsky-s-theory-of-language-learning/ https://www.scientificamerican.com/article/evidence-rebuts-c...
- bbor 3y agoThanks for posting, finally some support for his supposed debunking! Interesting reading for sure. That work fails to support Chomsky’s assertions. The research suggests a radically different view, in which learning of a child’s first language does not rely on an innate grammar module. Instead the new research shows that young children use various types of thinking that may not be specific to language at all—such as the ability to classify the world into categories (people or objects, for instance) and to understand the relations among things. These capabilities, coupled with a unique human ability to grasp what others intend to communicate, allow language to happen. The fact that very smart people think this refutes Chomsky makes me quite sad. They basically restated the UG theory in the last sentence, as proof that it’s wrong… Chomsky has been saying for literal decades that language is likely a corollary to the basic reasoning skills that set humans apart, but people still think UG means “kids are born knowing what a noun is” :(
- gizmo686 3y agoI'm reminded of a 400 leve linguistics class I took in undergrad. We had just read Chomsky's Remarks on Nominalization, and one of my classmates remarked, "I don't think this Chomsky guy understands X-Bar theory". The joke being that Chomsky was the major developer of X-Bar theory. We were just reading an early work of his. This also reminds me of evolution. Some people looked at discoveries in epigenetics and declared that it disproved Darwinian evolution in favor of Lamarckian evolution. Sure, Darwin's theory of natural selection combined with random variation at the point of reproduction does not explain 100% of evolution, but it is still covers most of it.
- csb6 3y ago
- effed3 3y agoI remember, some time ago, someone (more than one, to be precise) used NN (neuralnets) to resolve some math problems (differential equations, IIRC, i have lost the refs). Suddently some others stated the death of Mathematics, and in general, of Science as we all know, just pour all data in a NN and solution will appear, no more theory, models, hypothesis, experiments,.. needed. Nothing of this happened, but from time to time this -hallucination- return, missing the difference between Science and Technology, between a machine that work (or appear to work) and a model/theory that explain how/why.
- gwern 3y agoA better landing page (if you don't want to link the PDF) would be the official one: https://lingbuzz.net/lingbuzz/007180 https://lingbuzz.net/lingbuzz/007180
- dogcomplex 3y agoSavage, but maybe fair. If one of Chomsky's underlying claims is indeed that language requires innate hard parsing rules and can't just be derived from probabilistic sampling of a bunch of data - that seems completely dead in the water. It is entirely likely that the way we operate is probability-first, only deriving rules loosely after taking in lots of experiential data to speed up and simplify that initial fake-it-til-you-make-it understanding. The fact LLMs can get the quality we see using just this approach is a strong indicator that this method of understanding may be a fundamental approach of biological systems too. (and if you're arguing this is unfair because humans created the language that's being used for probabilistic training - well, look at image models trained on photographs instead and tell me those aren't an example of extreme quality derived purely from mass-inferenced data. Rules-based architectures don't necessarily need apply.) But honestly, this seems like a silly claim to begin with if it really was claimed. We have formal language theory complexity classes of probabilistic algorithms for a reason - they work! It shouldn't be surprising that the model can stretch down to the fundamentals too. Far fewer programmers (and linguists) were raised to think with these models than deterministic rules-based ones, but the field has been progressing alongside for decades, and now they get to play with powerful LLMs that take probabilistic inferencing to the extreme and will likely prove it works (very elegantly) for everything. This shouldn't be surprising in retrospect. Chomsky may very well be right that there always exists some fundamental elegant formula underlying any phenomenon (or at least any language). But it's undeniable at this point that simplistic statistical approaches can be applied at scale to those phenomenon and derive highly-useful general models, which also will very likely converge upon the elegant formulas he envisioned. The two are intrinsically linked, neither inseparable.