4 ms·
"We all saw the limitations of the old tools, and the benefits of the new" Probabilistic models work incredibly well, much better than transformational-generat
by hackandthink 4y ago
"We all saw the limitations of the old tools, and the benefits of the new"
Probabilistic models work incredibly well, much better than transformational-generative grammars.
- foobarqux 4y ago> Probabilistic models work incredibly well, much better than transformational-generative grammars. You've missed everything Chomsky said even though it is repeated in the article: Probabilistic models can be useful tools but they tell you nothing about the human language faculty (i.e. they are not science).
- visarga 4y agoThis kind of top-down approach misses the real hero - it's not the model, it's the data. 500GB of text can transform a randomly initialised neural net into a chatting, problem solving AI. And language turns babies into modern, functional adults. It doesn't matter how the model is implemented, they all learn more or less, the real hero is text. Let's talk about it more. It would have been interesting if Chomsky's approach could have predicted at what size of text we see the emergence of AI that passes the Turing test. Or even if it predicted that there is an emergent process in there.
- foobarqux 4y agoNo, again it's missing the point: None of this explains how the human language faculty works.
- morelisp 4y agoA child needs about 100KB of "text" a day to learn a language. If anything the data requirements of LLMs are proof positive they can't bear any relation to the human language faculty.
- pixl97 4y agoI mean we could be far overfeeding our data models too. Of course data models don't take years in real time to train either.
- morelisp 4y ago> I mean we could be far overfeeding our data models too. As far as I know evidence we have suggests the opposite, improvements are still mostly coming from more parameters and more data. > data models don't take years in real time to train either. Given the incommensurate architectures a fairer calculation of learning rate might be Wh - a human brain needs about 500Wh a day, GPT-3 was suspected to take about 1GWh to train.
- YeGoblynQueenne 4y agoI'm not well-informed on the subject, but I seem to remember that Chomsky's point was exactly on the data: his hypothesis about the human language faculty being innate (a "universal grammar", or "linguistic endowment" as he's been calling it more recently) was about the so-called "poverty of the stimulus". Meaning that human infants learn human languages while being exposed to pitiably insufficient amounts of data. Again, to my recollection, he based this on Mark E. Gold's result about language identification in the limit, which, simplifying, is that any language more complex than a finite language (in the Chomsky hierarchy of languages) is learnable only from an infinite number of positive examples, and languages more complex than regular languages also need an infinite number of negative examples. And those are labelled examples- labelled by an oracle. Since human language is usually considered to be at least context-free, and since infants are not exposed to infinite numbers of examples of their maternal languages, there must be some other element that allows them to learn such a language, and Chomsky called that a "universal grammar" etc. Still from memory, Chomsky's proposition also took account of data that showed that human parents do not give negative examples of language to their children, they only correct by giving positive examples (e.g. a parent would correct a child's grammar by saying something along the lines of "we don't say 'eated', we say 'eaten'"; so they would label a grammar rule learned by the child as incorrect -the rule that produced 'eated'- but they wouldn't give further negative examples of the same, or other rules, that produced similarly wrong instances, only a positive example of the correct rule. That's my interpretation anyway). Again all this is from memory, and probably half-digested. Wikipedia has an article on Gold's famous result: https://en.wikipedia.org/wiki/Language_identification_in_the_limit https://en.wikipedia.org/wiki/Language_identification_in_the... Incidentally, Gold's result, derived in the context of the field of Inductive Inference, a sort of precursor to modern machine learning, caused a revolution in machine learning itself. The very negative result caused Leslie Valiant to develop his PAC-Learning setting, that basically loosens the strong requirements for precision of Gold's identification in the limit, and so justified the focus of modern machine learning research to approximate, and efficient, learning. But that's another story.