2 ms·
Deep learning/neural networks -- as well as a number of other "machine learning" methods -- is fitting a mathematical model with a huge number of fitted (tunabl
by NumberSix 9y ago
Deep learning/neural networks -- as well as a number of other "machine learning" methods -- is fitting a mathematical model with a huge number of fitted (tunable) parameters and component functions to data. It has been known for a long time -- arguably back to the discovery of the Taylor series expansion in the early days of calculus -- that any function can be approximated arbitrarily well by the composition of an arbitrarily large composition of other functions.
If the task is interpolation between the data points this can be highly accurate. It the task is extrapolation such as prediction or designing a truly new machine or system in an engineering application, the approximation will often fail. The about one percent error rate in predictions of planetary motions from epicycles is one of the earliest cases of this problem.
The simple example is approximating data with a polynomial with an arbitrary number of terms such as in the Taylor series expansion. With enough terms a polynomial model can always approximate any data set arbitrarily well. Polynomial models can interpolate very well unless the data has some generally unusual behavior -- changing unpredictably at successively finer scales for example. However, finite polynomial models almost always extrapolate grossly incorrectly. As you move away from the data set in the space of independent variables such as X, the largest power N in X^N dominates and the polynomial approximation function blows up to either plus or minus infinity which is rarely physical.
What we think of as "understanding" corresponds conceptually in part to the ability to make accurate predictions. "Understanding" or "explanation" corresponds mathematically not to some arbitrary super-complex function with large numbers of arbitrary parameters but rather to a mathematical object such as as system of differential equations (e.g. Maxwell's Equations for Electromagnetism or the General Theory of Relativity) that express interrelationships among the variables and data points.
- ppod 9y agoMachine learning researchers are well aware of the problem of out-of-sample predictions, in fact most of the work goes into making systems appropriately generalizable by making architectures robust through dropout or regularization. The resulting systems might not correspond to our intuitive definition of "understanding", but they look more similar to what we know about the human brain than an elegant function with a small number of interpretable parameters.