3 ms·
I'm certainly convinced that sampling from a narrow dataset can lead to problems down the line. You don't remove possible bias in your data by randomly setting
by geebee 8y ago
I'm certainly convinced that sampling from a narrow dataset can lead to problems down the line. You don't remove possible bias in your data by randomly setting aside a testing set and using it for cross validation. It's an important practice, but yes, I agree that you can end up over training on structures that exist in both the testing and training set, since they're drawn from the same, narrow source.
As another example, you might train a model to identify positive and negative movie reviews (a pretty common example in intro tutorials). Your original data set might just be views by Roger Ebert. Your model, based on a training set, thinks it's at 100% accuracy. Cross validation on the test set reveals it's at 85%. Not bad! Then you apply it to reviews for 100 film critics. It's down to 70%. Then you apply it to random reviews left anonymously on the internet by 1000s of people. It drops all the way down to 55%.
That's a good argument in favor of a robust data set, drawn from multiple independent sources.
However, here's where I'm not convinced. How would a low training error based on a narrow data set be any less misleading than a low cross validation error based on a narrow training and testing set? If you aren't drawing from a robust data source, it seems the problem would be just as bad either way.
- sdenton4 8y agoVery true! But part of knowing that you don't know is recognizing when you're in unfamiliar territory. Being able to say "this example doesn't look like what I've seen before, and therefore I'm not as sure of my response." Algorithms which can express uncertainty /should be/ more robust to domain errors.