4 ms·
Yeah, figure 4 is more clear. This is the early stopped loss though, so the regularization is more explicit. If you trained the large models to completion on
by rprenger 6y ago
Yeah, figure 4 is more clear. This is the early stopped loss though, so the regularization is more explicit. If you trained the large models to completion on small data sets they would do much worse on test error due to overfitting.
It is interesting that larger models with regularization (early stopping) seems to work better than than training smaller models to convergence though.