5 ms·
I think that Norvig hits the nail on the head near the beginning of his piece: "I believe that Chomsky has no objection to this kind of statistical model [the
by mrow84 10y ago
I think that Norvig hits the nail on the head near the beginning of his piece:
"I believe that Chomsky has no objection to this kind of statistical model [the Newtonian model of gravitational attraction]. Rather, he seems to reserve his criticism for statistical models like Shannon's that have quadrillions of parameters, not just one or two."
This is no more than an objection to problems of fitting your chosen model to data. If you only have a small number of free parameters, then you can fit your model with a reasonable amount of data. If you have a large number of parameters then you have to introduce some extra assumptions, as Norvig (of course) acknowledges slightly earlier (described as "smoothing", in context):
"For example, a decade before Chomsky, Claude Shannon proposed probabilistic models of communication based on Markov chains of words. If you have a vocabulary of 100,000 words and a second-order Markov model in which the probability of a word depends on the previous two words, then you need a quadrillion (10^15) probability values to specify the model. The only feasible way to learn these 10^15 values is to gather statistics from data and introduce some smoothing method for the many cases where there is no data."
Thus, although both models are statistical, it is much easier to have confidence in Newton's law of gravitation than it is in a Markov model of some communication channel, because the data tell a clear picture. The imprecision of Newton's law in certain parts of the problem space (unobserved during his time) is a moot point - any such objections apply equally well to models with many parameters, and then you _still_ have to accept that you have made extra assumptions "outside" the scope of your model.
If you can explore your entire problem space, then you can build a complete "model". If not, then having more parameters than data _requires_ additional assumptions. Chomsky's point stands.
- grandalf 10y agoVery well articulated. One small point to add: Chomksy's interest has been in identifying the abstract characteristics of the language center of the human brain, which, for various reasons, does not seem likely to work like a Markov model. Analogously, one could look at the inputs and outputs of the human heart and potentially imagine a variety of physical structures that would explain them, and some of those physical structures would be biologically real/plausible and others would not. Things like constraints on working memory, exposure to input, and cross-language studies have informed the constraints that Chomsky has proposed to determine what kind of model would best capture the essential quality of the brain system. link: https://www.amazon.com/Knowledge-Language-Nature-Origins-Convergence/dp/0275917614 https://www.amazon.com/Knowledge-Language-Nature-Origins-Con...
- foobarqux 10y agoNo, Chomsky's problem with statistical models is that you can't learn much about the underlying system from them, and so it isn't science, which has as a fundamental aim to understand the world. Statistical models can be useful but they generally aren't meaningful to scientific progress.
- dragonwriter 10y ago> No, Chomsky's problem with statistical models is that you can't learn much about the underlying system from them, and so it isn't science, which has as a fundamental aim to understand the world. Science has the fundamental aim to make testable predictive models of the world; any "understanding" other than that represented by a testable predictive model is irrelevant to science, except insofar as it might provide intuition on which to found hypotheses of better predictive models. Statistical models, to the extent that they are testable and predictive, are exactly the kind of thing that science is about.
- foobarqux 10y agoPutting aside whether how we want to define "science" there are models which attempt to represent the underlying system (e.g. generative grammars) and other "models" which merely try to predict the input/output behavior of those systems (e.g. neural nets). The latter typically doesn't lend any understanding of the underlying system (e.g. we don't understand much about language from a deep neural net translation system), which isn't surprising because they are designed for prediction not for representing the internal behavior of the system. To use one of Chomsky's examples, if you built a deep neural net that could predict the weather with 100% accuracy then you have performed an amazing feat of engineering that is incredibly useful to the world. But you haven't necessarily learned anything about the weather.
- dragonwriter 10y agoThe problem there seems to be not one with statistical models, but with the fact that extracting a parisomonious mathematical model from a neural net is nontrivial (presumably, you can extract a mathematical model from the internal weights, etc., but for even known simple relationships neural net models produced without an a priori prediction of the general shape of the relationship tend to be very complex.) Of course, were the model really 100% accurate, then converting it to a the most parsimonious perfectly accurate model would be simply a matter of reduction; the real problem is that real neural net models are not 100% accurate, and are often far more complex than more accurate models, and often aren't convenient to analyze to deduce the more-accurate and simpler model.