4 ms·
> They test against XGBoost, GBDT Auto-sklearn, and others. Did you read the paper? Yes. Did you read my comment? They compare NN + Cocktail vs. vanilla XGB.
by nightcracker 5y ago
> They test against XGBoost, GBDT Auto-sklearn, and others. Did you read the paper?
Yes. Did you read my comment?
They compare NN + Cocktail vs. vanilla XGB. They don't compare NN + Cocktail vs. XGB + Cocktail.
To make it crystal clear, if I wrote a paper "existing medicine A enhanced with novel method B is more effective than existing medicine C" and I did not include the control "C + B" (assuming if relevant, which is the case here), that'd be bad science. It's very much possible that novel method B is doing the heavy lifting and A isn't all that relevant. s/A/NN, s/B/Cocktail, s/C/XGBoost.
- ipsum2 5y agoHow would you even apply layer normalization or SWA to XGB? These methods are neural net specific.
- nightcracker 5y agoBatch normalization is nothing neural network specific to it if you use it on the input layer. I don't think it matters for a tree algorithm like XGBoost either way though. SWA is pretty NN specific. So leave it out for XGB. There's a bunch that are relevant, and they could be very important.
- jackylupino 5y agoBatch norm has an advantage for iterative methods on mini-batches, while XGB uses the full training set. Using batch norm on the full training set is equivalent to Z-normalizing the features, which has no effect at all for XGB as the scale of features plays no role at the split decisions of the tree nodes. Apart few non-parametric data augmentation methods (notice adversarial augmentation is also nn specific), I do not think any other regularization used in that paper can be directly/intuitively applied to XGB.
- mandelken 5y agoGBDT have their own set of hyperparameters such as learning rate, number of trees, min samples per bin, l0, l1, etc. So you could definitely also create an appropriate cocktail to optimize on, although GBDT are typically more robust wrt huperparameters.
- nightcracker 5y agoThe authors do claim to do a hyperparameter sweep but only for vanilla XGB hyperparams.
- civilized 5y agoThe old "my method (with as much optimization as I could get away with) beats the other method (with as little optimization as I could get away with)"
- disgruntledphd2 5y agoYup. The Least Publishable Unit strikes again.
- blt 5y agoI don't think every single regularization method in the cocktail can be applied to non-neural-network methods, but I'm pretty sure some of them can, like data augmentation. The authors could have figured out which methods can be applied to non-NN models or considered if equivalent/analogous methods exist. I agree it would make a fairer comparison.
- jackylupino 5y agoI agree with your point, e.g. data augmentation can be added, but thats pretty much it. All the other regularization techniques they use are neural network specific and cannot be applied to gradient-boosted trees. What I find particularly striking at this paper is that their method trains a single neural network which outperforms an ensemble of decision trees (XGBoost). Asking for perfect apple-to-apple comparisons means also comparing an ensemble of the MLPs vs. XGBoost. In this context, at least the message here is that XGBoost and/or other gradient-boosted methods are not anymore a silver bullet for tabular datasets. Boosting for trees was great in reducing both bias and variance, but apparently neural networks can achieve the same effect with a high capacity (low bias) and a mix of modern regularization techniques (low variance).