3 ms·
I would take the opposite approach here. Instead of overfitting, bake probability into your model. Whether using a Bayesian weights or using an ad-hoc version (
by workingon 4y ago
I would take the opposite approach here. Instead of overfitting, bake probability into your model. Whether using a Bayesian weights or using an ad-hoc version (I.e. dropout and batch normalization), you can do your predictions in ensemble and look at the deviation of predictions. the combination of this method and a small dataset usually leads to wider ranges of prediction, which can be interpreted as a model uncertainty. When using these techniques we have found smaller datasets lead to more uncertainty, but it often will bound the error. I think this is much more useful than assuming a small dataset means an overfit model will do good on future data.