5 ms·
The author mentions he defines over fit as “Test error is always larger than training error”. Is there an algorithm or model where that’s not the case?
by syntaxing 3y ago
The author mentions he defines over fit as “Test error is always larger than training error”. Is there an algorithm or model where that’s not the case?
- jgalt212 3y agoYeah, that's a crummy definition. You can easily force "Test error is always larger than training error" for any model type.
- deleted 3y ago[deleted]
- jphoward 3y agoYou can see regularly in practice where aggressive data augmentation is used, which obviously is only used on training data. But, of course, you'd still be 'overfit' if you fed in unaugmented training data.
- jksk61 3y agoYes and no. Suppose for example you give MNIST to a SVM and fit the model, then test it only on 0 and 1 digits, which are generally well discriminized, you'll get almost 100% accuracy in test, whereas 97% or less in training. (ye probably need some preprocessing, like using PacMAP or UMAP or whatever but the point is the same) However, that's just because I decided the right data to test it onto. So, you can't really say much on a model using that definition.
- dataflow 3y agoI think they mean "always (statistically) significantly larger". They're probably imagining that something like cross-validation would make test errors approximately equal to training errors, but if you consistently see significantly larger errors, then you've overfit.
- bbstats 3y agoWhen you find the minima of your validation curve, rarely if ever is your test loss lower than your training loss. I don't think this necessarily means you're overfitting.
- onos 3y agoA pedantic example: a model that ignored the training data would do just as well on the training set as on the test set.
- tesdinger 3y agoIf you do not train on the training set, then there is no training set, and your example is degenerate.
- maxbond 3y agoWe've trained on the training set, but our fit() implementation was, "accept the input, do nothing with it, and then return." When we then evaluate the model on the test set and training set, assuming there isn't a large covariant shift between the two or something, we'll get approximately equal values. So it's a degenerate case, but not of the training set. (And presumably that's partly what they meant by "pedantic".)
- PartiallyTyped 3y agoUnder the Probably Approximately Correct framework, there's an x% chance your error on the empirical distribution deviates more than \epsilon away in either direction from that of the actual distribution, so one can assume similar things between two empirical distributions.