3 ms·
If you don't get the holdout set by shuffling, how do you get it?
by vlmutolo 5y ago
If you don't get the holdout set by shuffling, how do you get it?
- datastoat 5y agoIdeally, there's some field in our data that we can use -- e.g. train on data from cities in several countries, test on data from cities in a different country; train on one corpus, test on another; train on climate data at one level of CO2, test on data at another. If there's no such field, I do a simple 1d dimension reduction, and pick off 20% of the data at the extremes. If we know exactly the domain where our model will be used, and it matches exactly the disposition of the training data, then shuffling is fine. But if we want our model to work well in new domains, then we're like scientists looking for generalizable laws, and the best way to test this is by testing in unseen domains. It's often hard to get hold of data with this domain-diversity, which is why a lot of ML models are fragile.
- monkeybutton 5y agoWalk the walk! If your data has any notion of temporality, train on [t-K, t] and predict t+1.
- PeterisP 5y agoAs you often want to predict generalization to future unseen data, it's important to consider how the unseen data will be different from your current dataset - all the theory operates on the simplifying assumption that your dataset is an IID random sample from the "true" data but, of course, usually that's not really the case and you know and expect that the future data will have a different distribution than what you have - economics and politics trends will have new major events; future textual data will see new names, event names, terms and concepts that did not exist in 2020; etc, etc. So if you want to estimate actual generalization, you have to at least try to isolate some of these aspects. For text, if you take a random sample of shuffled sentences then you'll get a different outcome than if you shuffle whole documents, because it'll be "cheating" as your test sentences will always have some relevant context in your training set, which won't be the case for real unseen data which will sometimes introduce totally new things. And if you train something on data gathered 2015-2020 and evaluate on data from 2021, then you'll likely get worse (but more informative!) measurements than training on a random sample chosen from the whole 2015-2021 range, simply because there are major world events and 'distribution shifts' over time. IMHO a true test for generalization of many approaches would be to train stuff on data before 2020 and test on data from 2020 - to see how well it generalizes given a major event like Covid pandemic that changes all kinds of aspects everywhere in society that generates the new data.
- wodenokoto 5y agoFor time series data, you don't want to shuffle, because now you are predicting the past based on the future. I think the best way is to choose a minimum timespan you want to be able to predict on, train on that and predict on the future, then retrain, including that future and predict on the next future. If your time series allow it (e.g. contains a bunch of independent target groups) you can train/test on a subset and leave a another subset for validation