3 ms·
Ideally, there's some field in our data that we can use -- e.g. train on data from cities in several countries, test on data from cities in a different country;
by datastoat 5y ago
Ideally, there's some field in our data that we can use -- e.g. train on data from cities in several countries, test on data from cities in a different country; train on one corpus, test on another; train on climate data at one level of CO2, test on data at another. If there's no such field, I do a simple 1d dimension reduction, and pick off 20% of the data at the extremes.
If we know exactly the domain where our model will be used, and it matches exactly the disposition of the training data, then shuffling is fine. But if we want our model to work well in new domains, then we're like scientists looking for generalizable laws, and the best way to test this is by testing in unseen domains. It's often hard to get hold of data with this domain-diversity, which is why a lot of ML models are fragile.