4 ms·
RF's work well on heterogenous data (ie a mix of data types from difrent sources which might include numerical, categorical and text data etc. They also tend to
by micro_cam 11y ago
RF's work well on heterogenous data (ie a mix of data types from difrent sources which might include numerical, categorical and text data etc. They also tend to be fairly resistance to noise in the data and to work well on data sets that are wide (ie have more features/dimensions then cases/observations). Finally they are fairly resistant to overfitting (ie fitting quirks in the training data instead of generalizable signal) and don't need too much parameter tuning.
This is largely because they are a randomized ensemble of weaker models. Individual decision trees are quite prone to overfitting and other issues but in an rf you grow a bunch of them on diffrent bootstrap samples of the data and let them vote and it turns out the combined performance is much better and much less error prone then a single model.
Specific examples where they work well include genetic data (many more noisy variables then observations) and customer/consumer data. They also get used in image data and signal processing but deep neural networks are recently tending to beat them here and in similar less heterogenous data sets.