3 ms·
Random Forests are nice in having few parameters-- number of trees, number of restricted features to sample on at each decision node. It's the variations on th
by textminer 13y ago
Random Forests are nice in having few parameters-- number of trees, number of restricted features to sample on at each decision node.
It's the variations on the standard RF that have interested me. Rotation Forests, where features are partitioned randomly and rotated (by PCA or random projection) before decision boundaries are drawn. Extremely Randomized Forests, where the node splits are completely random and not based on best possible Gini/Entropy gain along some feature. There's even an interesting use of deterministic annealing out there for incorporating unlabeled data points in an attempt at semisupervised learning.
These each have their own parameters to tune, but have had slightly different performance on different problem domains. Even models require different data-- most decision trees have a quite natural way of imputing missing values, but something like a Rotation Forest can handle neither missing values nor categorical data (unless you map it to m binary features). And that complexity spooks me away from Machine Learning as a Service, where one could start failing to understand his or her models. (Plus then I'd probably be out of a job.)