13 ms·
>It is not true that a piecewise-linear model trained on a set of data points will produce only outputs which are a convex combination of outputs that appear in
by yldedly 5y ago
>It is not true that a piecewise-linear model trained on a set of data points will produce only outputs which are a convex combination of outputs that appear in the training set.
No, and I didn't claim that. I said that, outside the training sample, the model is linear (or quadratic in the case of transformers, thanks for pointing that out) Whether linear or quadratic, a model that has a fixed structure outside the training sample, will obviously not fit data which lies far away from the training sample - i.e. it will not extrapolate. This isn't controversial - it's just something people like to forget about.
>A model trained on images which produced only convex combinations of images in its training set, would clearly be producing what could be called “interpolations between images in its training set”, and taking convex combinations of images is unimpressive.
True! I should have clarified that it's not linear interpolation in pixel space (or input space generally), but interpolation on the latent manifold. This is where the power, as well as limitations of deep learning come from.
It's definitely non-trivial to identify the latent manifold of data - different dimensions of the manifold may sometimes even correspond to independent components, as you mention (position of eyes, skin tone,...) (though empirically, finding disentangled latent codes is mostly a function of the random seed).
How does an NN process a new input? It maps the input to the latent manifold.
In the input space, it will be some highly non-linear, non-trivial combination of points, which in terms of Euclidean distance in the input space, could be arbitrarily close or far away.
In the latent space, the output will be some convex combination of nearby points.
Here's the kicker - even if your problem happens to be well-modeled as a continuous, low-dimensional manifold embedded in a high-dimensional space (and many, many problems aren't), and even if you manage to obtain a super dense sampling of input space, so that the manifold can be well-approximated (which is impractical or impossible for most problems),
you will never be able to generalize beyond the data distribution.
Our brains don't stop working as soon as conditions are slightly different from what we've seen before. If there's a slight fog on a Stop sign, we can still see a stop sign. If the Go board is 9x9 rather than 19x19, we can still play Go. If we can play Starcraft on one map, we're pretty much as good on a different map, we don't need to relearn the game over the next several thousand years.
How come? Because we aren't just latent space interpolators. We can extrapolate.