3 ms·
The article is good, but "The dangers of cross-validation" section is wrong. Cross-validation means you split your data into folds and run multiple experiments
by kmike84 9y ago
The article is good, but "The dangers of cross-validation" section is wrong.
Cross-validation means you split your data into folds and run multiple experiments, getting better data efficiency as compared to a single split. Splits don't have to be random.
scikit-learn provides GroupKFold and TimeSeriesSplit objects to use with cross-validation which address exactly the problems described in the article.
Cross-validation is not a panacea, in real world there are many other issues - e.g. it is common to have a training dataset which doesn't represent production data distribution, often because such dataset turns out to be much cheaper to get.
Andrew Ng's http://www.mlyearning.org/ http://www.mlyearning.org/ is a great resource on this; it discusses real-world problems with train/test/validation splits and gives practical advice, can't recommend it highly enough.