3 ms·
I haven't read much on gradient boosting, so questions: 1. Where is the gradient? This explanation makes it sound like a straight Generalized Additive Model.
by ced 10y ago
I haven't read much on gradient boosting, so questions:
1. Where is the gradient? This explanation makes it sound like a straight Generalized Additive Model.
2. In fact, the explanation makes it sounds worse than random forests. Wouldn't it quickly overfit? Where does the boosting come into play?
- sforzando 10y agoThese helpful, well-written slides help explain where the "gradient" comes into "Gradient Boosting": http://www.ccs.neu.edu/home/vip/teach/MLcourse/4_boosting/slides/gradient_boosting.pdf http://www.ccs.neu.edu/home/vip/teach/MLcourse/4_boosting/sl... The gist of it is: when you add a new decision tree that fits to the residual error, this new tree is fitting to the negative gradient of the loss function (ie training error). Thus, adding the new decision tree to your existing ensemble takes a gradient-descent step that seeks to minimize the loss function (ie training error). Boosting comes in because the model is combining several weak learners/models (individual trees) into a strong learner (ensemble of trees). Each individual tree breaks up the input space into piecewise-constant regions that best approximate the target function. This representation will incur some error - thus, a new tree is fit to minimize the error over the entire input space, ie by breaking up the input space into piecewise-constant regions, etc. So, it's boosting not in the traditional Adaboost sense: where the final model is a linear combination of "dumb" classifiers. Instead, I'd liken it more to a cascade method: each tree T_{n} seeks to fix the errors from the previous tree T_{n-1}: https://en.wikipedia.org/wiki/Cascading_classifiers https://en.wikipedia.org/wiki/Cascading_classifiers There's actually a cool facial landmark detector that uses this same cascading idea to train an extremely fast (and quite accurate) system. In essence, they use a cascade of random forests (in a gradient-boosting framework) to detect landmarks. The dlib library has a great implementation, along with a pretrained model. I've used it in my research, and while not perfect, have been satisfied with its results: http://blog.dlib.net/2014/08/real-time-face-pose-estimation.html http://blog.dlib.net/2014/08/real-time-face-pose-estimation.... http://www.cv-foundation.org/openaccess/content_cvpr_2014/papers/Kazemi_One_Millisecond_Face_2014_CVPR_paper.pdf http://www.cv-foundation.org/openaccess/content_cvpr_2014/pa...
- ced 10y agoThose slides were very helpful, thank you.
- deleted 10y ago[deleted]
- selectron 10y agoThe explanation glosses over a few important details. Gradient boosting works by adding some small weight to the instances the model is incorrectly predicting. The amount of extra weight these instances get is a parameter that is tuned with validation - because this parameter can be 0, if you are doing correct cv gradient boosting trees is usually superior to random forests. You also do need to tune the number of trees you use in gradient boosting or else you will overfit. Gradient boosting doesn't get nearly enough hype as compared to things like neural nets. The significant majority of winning solutions to Kaggle competitions for a non-image or text-processing dataset will use xgboost to do gradient boosting as part of the ensemble model. Furthermore, it is a really easy method to understand and use while still being state-of-the-art.