4 ms·
> Results show that tree-based models remain state-of-the-art on medium-sized data (∼10K samples) even without accounting for their superior speed. Is that rea
by Permit 4y ago
> Results show that tree-based models remain state-of-the-art on medium-sized data (∼10K samples) even without accounting for their superior speed.
Is that really "medium"? That seems very small to me. MNIST has 60,000 samples and ImageNet has millions.
I think the title overstates the findings. I'd be interested to hear how these methods compare on much larger datasets. Is there a threshold at which deep learning outperforms tree-based models?
Edit: They touch on this in the appendix:
> A.2.2 Large-sized datasets
> We extend our benchmark to large-scale datasets: in Figures 9, 10, 11 and 12, we compare the results
of our models on the same set of datasets, in large-size (train set truncated to 50,000 samples) and
medium-size (train set truncated to 10,000 samples) settings.
> We only keep datasets with more than 50,000 samples and restrict the train set size to 50,000 samples
(vs 10,000 samples for the medium-sized benchmark). Unfortunately, this excludes a lot of datasets,
which makes the comparison less clear. However, it seems that, in most cases, increasing the train set
size reduces the gap between neural networks and tree-based models. We leave a rigorous study of
this trend to future work.
- mochomocha 4y agoI've put in production numerous models with millions of tabular data points and a 10^5-10^6 feature space where tree-based models (or FF nets) outperform more complex DL approaches.
- whymauri 4y agoI'll chime in with billions of data points and 100-300 feature space with some smart feature engineering outperforming DL in runtime/compute (by orders of magnitude) and performance. But the domain was very specific and everything prior to the tree was doable with matrix operations, with the tree model summarizing a mixture of experts that chose optimal features.
- bloudermilk 4y agoDamn, this sounds fascinating! Have you shared more anywhere eelse?
- adamsmith143 4y agoWhat kind of freakish tabular data do you have with a million columns??
- ramraj07 4y agoA badly defined one probably? At least for one of the test arms.
- mochomocha 4y agoCategorical variables in the datasets of large tech companies can take a lot of different values.
- beckingz 4y agoMany real world problems that result in data are decidedly medium: small enough to fit in excel, large enough to be too big to comfortable handle in excel.
- ramraj07 4y agoThen these real world problems don’t actually warrant deep learning ? I thought the biggest leap in NN and deep learning in recent history was the realization that we need a ton of data to get maximal effectiveness from them; it now sounds counterproductive to forget this and cry they don’t work well with 10,000 rows.
- aqsalose 4y ago> Then these real world problems don’t actually warrant deep learning ? It is an important lesson to be communicated. I'd like to present a conjecture: everyone thinks their data is big until they have worked on much larger dataset. ("We have 10k samples, it is quite big!" -> "We have 1m data records, is quite big!" -> "Our process outputs that much per day")
- beckingz 4y agoThat's right. Most problems don't warrant deep learning with the current state of the art.
- riedel 4y agoMNIST is not your typical real world tabular data. Many if not most data science problems out there are still in the range of a few k samples from my perspective (trying to "sell" ML to the average company) From a statistical point of view I would not call the datasets small (you can decently compare two means from subsets without needing a student's distribution).
- nonameiguess 4y agoAssuming the categories are meant to apply to any data sets, anything amenable to machine learning at all is at least medium data. "Small" data would be something like a human trial with n=6 because the length and compliance of the protocol is so onerous. There are entirely different statistical techniques for finding significance in the face of extremely low power.