3 ms·
I'm going off a first pass through the paper, but it appears that what this paper shows is that the training error can be 0 on an entirely randomized data set,
by geebee 8y ago
I'm going off a first pass through the paper, but it appears that what this paper shows is that the training error can be 0 on an entirely randomized data set, but the generalization error - the difference between the error on the test set and the training set, does increase dramatically as label corruption increases.
My understanding is that cross validation does multiple combinations of splitting the input data into test and training sets... so if cross validation measures the generalization error, wouldn't this catch the low predictive value resulting from randomization of labels or input?
I'm not saying the paper doesn't have value, but I think it's more about the fact that neural nets can obtain a training error of zero on randomized data, not a testing error (or generalization error, which represents the difference between training error and testing error, as far as I can tell).
To be clear, I'm not an expert, and this is just what I gleaned from a first pass over the paper.
- sdenton4 8y agoAll true. The interesting thing here is that the neural network has /no idea/ that it sucks at generalization, though. Yes, we can do extra work to calibrate outputs, but it would be much better to have some idea of uncertainty from the network itself. (Added as edit) also keep in mind that datasets themselves often fail to generalize - overriding to a particular set makes for domain error when moving to slightly different data. Cross validation won't help wit that, but more "self aware" algorithms might.
- geebee 8y agoBut... isn't that the entire point of splitting your initial training data into a training set and a separate testing set? Why is it better to have an idea of uncertainty from the model itself when you can get the generalization error through cross validation, or by setting aside a testing set? It's interesting to see that a neural net will reach a training error of zero on randomized data, and it's a worthwhile contribution to the literature to demonstrate this, test it, and measure it... but the outcome here doesn't surprise me. From experience I know that random forests will also show nearly 100% accuracy on a training set but show far lower accuracy for a testing set, so while I think it's great to measure it, the conclusion in this paper is not surprising. In no way is that a knock on the paper, people weren't surprised that Fermat's last theorem turned out to be true, but that doesn't make the proof any less of an accomplishment!
- sdenton4 8y agoFirst, because the dataset itself might be biased in subtle ways, in which case your cross validation won't help. This happens All. The. Time. For example, your training set for speech recognition might use a nice microphone uniformly, and everything goes to hell once you deploy to cell phones, because the microphones have different characteristics. Or, in the case of financial markets, the future might not look like the past. And the present might not look like the past... So you get datasets that are very time-specific, and thus prone to overfitting to local conditions and/or noise. Secondly, you can absolutely overfit your cross validation set, same as p-hacking. Run experiments until you have a slight positive, statistically significant result, then tell your managers you've got some crazy new sliver of alpha. And then when it hits really new data, it falls to pieces, because repeated experiments on noise will eventually produce a statistically significant result. It's like the old saw about freshmen who don't know that they don't know... Our current ML models tend to be like freshmen, or freshmen with bandaids...
- geebee 8y agoI'm certainly convinced that sampling from a narrow dataset can lead to problems down the line. You don't remove possible bias in your data by randomly setting aside a testing set and using it for cross validation. It's an important practice, but yes, I agree that you can end up over training on structures that exist in both the testing and training set, since they're drawn from the same, narrow source. As another example, you might train a model to identify positive and negative movie reviews (a pretty common example in intro tutorials). Your original data set might just be views by Roger Ebert. Your model, based on a training set, thinks it's at 100% accuracy. Cross validation on the test set reveals it's at 85%. Not bad! Then you apply it to reviews for 100 film critics. It's down to 70%. Then you apply it to random reviews left anonymously on the internet by 1000s of people. It drops all the way down to 55%. That's a good argument in favor of a robust data set, drawn from multiple independent sources. However, here's where I'm not convinced. How would a low training error based on a narrow data set be any less misleading than a low cross validation error based on a narrow training and testing set? If you aren't drawing from a robust data source, it seems the problem would be just as bad either way.