4 ms·
Paper is missing the control: how good is this 'cocktail of regularization' when applied to traditional methods like XGBoost? At best you can claim the result
by nightcracker 5y ago
Paper is missing the control: how good is this 'cocktail of regularization' when applied to traditional methods like XGBoost?
At best you can claim the result here that neural networks with regularization methods can beat traditional methods without it, but to be apples to apples both methods must have access to the same 'cocktail of regularization'.
- audi0slave 5y agoThis paper compares xgboost vs nueural nets and an ensemble [Tabular Data: Deep Learning is Not All You Need] https://arxiv.org/abs/2106.03253 https://arxiv.org/abs/2106.03253
- ipsum2 5y agoFrom the paper: > This paper is the first to provide compelling evidence that wellregularized neural networks (even simple MLPs!) indeed surpass the current state-of-the-art models in tabular datasets, including recent neural network architectures and GBDT (Section 6). > Next, we analyze the empirical significance of our well-regularized MLPs against the GBDT implementations in Figure 2b. The results show that our MLPs outperform both GBDT variants (XGBoost and auto-sklearn) with a statistically significant margin. They test against XGBoost, GBDT Auto-sklearn, and others. Did you read the paper?
- nightcracker 5y ago> They test against XGBoost, GBDT Auto-sklearn, and others. Did you read the paper? Yes. Did you read my comment? They compare NN + Cocktail vs. vanilla XGB. They don't compare NN + Cocktail vs. XGB + Cocktail. To make it crystal clear, if I wrote a paper "existing medicine A enhanced with novel method B is more effective than existing medicine C" and I did not include the control "C + B" (assuming if relevant, which is the case here), that'd be bad science. It's very much possible that novel method B is doing the heavy lifting and A isn't all that relevant. s/A/NN, s/B/Cocktail, s/C/XGBoost.
- ipsum2 5y agoHow would you even apply layer normalization or SWA to XGB? These methods are neural net specific.
- nightcracker 5y agoBatch normalization is nothing neural network specific to it if you use it on the input layer. I don't think it matters for a tree algorithm like XGBoost either way though. SWA is pretty NN specific. So leave it out for XGB. There's a bunch that are relevant, and they could be very important.
- jackylupino 5y agoBatch norm has an advantage for iterative methods on mini-batches, while XGB uses the full training set. Using batch norm on the full training set is equivalent to Z-normalizing the features, which has no effect at all for XGB as the scale of features plays no role at the split decisions of the tree nodes. Apart few non-parametric data augmentation methods (notice adversarial augmentation is also nn specific), I do not think any other regularization used in that paper can be directly/intuitively applied to XGB.
- mandelken 5y agoGBDT have their own set of hyperparameters such as learning rate, number of trees, min samples per bin, l0, l1, etc. So you could definitely also create an appropriate cocktail to optimize on, although GBDT are typically more robust wrt huperparameters.
- nightcracker 5y agoThe authors do claim to do a hyperparameter sweep but only for vanilla XGB hyperparams.
- civilized 5y agoThe old "my method (with as much optimization as I could get away with) beats the other method (with as little optimization as I could get away with)"