3 ms·
If the labels are random they don't correspond to anything, so any "features" are essentially noise. The classifier is, in effect, memorising every element in
by robert_tweed 9y ago
If the labels are random they don't correspond to anything, so any "features" are essentially noise.
The classifier is, in effect, memorising every element in the training set. It's training a compression algorithm for storing that data.
It should be noted that "training a compression algorithm" isn't always a bad thing per se, because that's how autoencoders work, which is one of the main ways to do deep learning.
The key term in the article is "the effective capacity" of the model. If you have a big enough network, it can simply memorise everything you give it. This makes it difficult to know if such a model will generalise. A much smaller model won't overfit in the same way, but also might not perform as well as a larger, more sophisticated model. The problem in deep learning is nobody can tell how much of the training data has simply been saved somewhere in the model (in an obfuscated and compressed way).
There is some related research about reconstructing the training data from deep networks (which has privacy implications), but I don't have a link handy.