3 ms·
Nice post. However, I was kind of surprised to see the author list interpretability as one of the drawbacks of random forests: > Another point that some might
by atpaino 11y ago
Nice post. However, I was kind of surprised to see the author list interpretability as one of the drawbacks of random forests:
> Another point that some might find a concern is that random forest models are black boxes that are very hard to interpret.
I generally agree that random forests are more difficult to interpret than linear models such as logistic regression, but I think they're still far more interpretable than more comparable non-linear models such as neural networks or SVMs with non-linear kernels. At the end of the day, a random forest is just a bunch of decision trees, each of which are very straightforward for humans to understand. Additionally, there are a number of straightforward methods available for assessing the importance of each individual feature in a random forest in aggregate ([1], section 15.4). Neural networks, on the other hand, result in models that are too cryptic to be meaningfully inspected by a human, and have more complex variable importance measures [2].
The relative interpretability of these non-linear models played a large factor in our decision to add random forests to our modeling stack at Sift Science, which you can read more about here: http://blog.siftscience.com/blog/2015/large-scale-decision-forests-lessons-learned http://blog.siftscience.com/blog/2015/large-scale-decision-f...
[1]: http://statweb.stanford.edu/~tibs/ElemStatLearn/ http://statweb.stanford.edu/~tibs/ElemStatLearn/
[2]: http://www.massey.ac.nz/~mkjoy/pdf/Olden,Joy&DeathEM.pdf http://www.massey.ac.nz/~mkjoy/pdf/Olden,Joy&DeathEM.pdf