4 ms·
I've implemented a random forest library that achieves similar or better results to some of the ones described [1] and in my experience there is a lot of room f
by micro_cam 11y ago
I've implemented a random forest library that achieves similar or better results to some of the ones described [1] and in my experience there is a lot of room for interpretation in the algorithm. In most of these cases there isn't necessarily a right way, one implementation may handle clean data better while another handles dirty/noisy data better.
Some examples of spots where interpretations differ are:
1) Do you exhaustively or randomly search for categorical splits, require one hot encoding or use some other heuristic?
2) Do you use 32 or 64 bit floats? Or do you support both but stop tree growth but have a minimum impurity decrease constant to make sure tests pass on both (scikit does this, it is essentially an unexposed hyper parameter).
3) How do you handle features that have become constant in the branch of a tree being extended? Scikit insists that at least one non constant feature be examined for each potential split resulting in bigger trees, other implementations will examine only m features even if they are all constant. I made this switchable in mine, scikit behavior works better on many test data sets, the other way on some noisy overfitting prone datasets.
4) If default paramaters are used do you default to gini impurity or entropy? Do you round to the nearest number or just up (so you always get at least one) after taking the square root of the number of features to set the number of features examined for each split.
All of this just means it is important to try different implementations and do hyperparameter searches etc on the data you actually care about.
[1] https://github.com/ryanbressler/CloudForest https://github.com/ryanbressler/CloudForest