3 ms·
Batch norm has an advantage for iterative methods on mini-batches, while XGB uses the full training set. Using batch norm on the full training set is equivalent
by jackylupino 5y ago
Batch norm has an advantage for iterative methods on mini-batches, while XGB uses the full training set. Using batch norm on the full training set is equivalent to Z-normalizing the features, which has no effect at all for XGB as the scale of features plays no role at the split decisions of the tree nodes. Apart few non-parametric data augmentation methods (notice adversarial augmentation is also nn specific), I do not think any other regularization used in that paper can be directly/intuitively applied to XGB.