4 ms·
But... isn't that the entire point of splitting your initial training data into a training set and a separate testing set? Why is it better to have an idea of u
by geebee 8y ago
But... isn't that the entire point of splitting your initial training data into a training set and a separate testing set? Why is it better to have an idea of uncertainty from the model itself when you can get the generalization error through cross validation, or by setting aside a testing set?
It's interesting to see that a neural net will reach a training error of zero on randomized data, and it's a worthwhile contribution to the literature to demonstrate this, test it, and measure it... but the outcome here doesn't surprise me. From experience I know that random forests will also show nearly 100% accuracy on a training set but show far lower accuracy for a testing set, so while I think it's great to measure it, the conclusion in this paper is not surprising.
In no way is that a knock on the paper, people weren't surprised that Fermat's last theorem turned out to be true, but that doesn't make the proof any less of an accomplishment!
- sdenton4 8y agoFirst, because the dataset itself might be biased in subtle ways, in which case your cross validation won't help. This happens All. The. Time. For example, your training set for speech recognition might use a nice microphone uniformly, and everything goes to hell once you deploy to cell phones, because the microphones have different characteristics. Or, in the case of financial markets, the future might not look like the past. And the present might not look like the past... So you get datasets that are very time-specific, and thus prone to overfitting to local conditions and/or noise. Secondly, you can absolutely overfit your cross validation set, same as p-hacking. Run experiments until you have a slight positive, statistically significant result, then tell your managers you've got some crazy new sliver of alpha. And then when it hits really new data, it falls to pieces, because repeated experiments on noise will eventually produce a statistically significant result. It's like the old saw about freshmen who don't know that they don't know... Our current ML models tend to be like freshmen, or freshmen with bandaids...
- geebee 8y agoI'm certainly convinced that sampling from a narrow dataset can lead to problems down the line. You don't remove possible bias in your data by randomly setting aside a testing set and using it for cross validation. It's an important practice, but yes, I agree that you can end up over training on structures that exist in both the testing and training set, since they're drawn from the same, narrow source. As another example, you might train a model to identify positive and negative movie reviews (a pretty common example in intro tutorials). Your original data set might just be views by Roger Ebert. Your model, based on a training set, thinks it's at 100% accuracy. Cross validation on the test set reveals it's at 85%. Not bad! Then you apply it to reviews for 100 film critics. It's down to 70%. Then you apply it to random reviews left anonymously on the internet by 1000s of people. It drops all the way down to 55%. That's a good argument in favor of a robust data set, drawn from multiple independent sources. However, here's where I'm not convinced. How would a low training error based on a narrow data set be any less misleading than a low cross validation error based on a narrow training and testing set? If you aren't drawing from a robust data source, it seems the problem would be just as bad either way.
- sdenton4 8y agoVery true! But part of knowing that you don't know is recognizing when you're in unfamiliar territory. Being able to say "this example doesn't look like what I've seen before, and therefore I'm not as sure of my response." Algorithms which can express uncertainty /should be/ more robust to domain errors.