8 ms·
I have a theory - tree based models require minimal feature engineering. They are capable of handling categorical data in principled ways, they can handle the m
by CapmCrackaWaka 4y ago
I have a theory - tree based models require minimal feature engineering. They are capable of handling categorical data in principled ways, they can handle the most skewed/multimodal/heteroskedastic continuous numeric data just as easily as a 0-1 scaled normal distribution, and they are easy to regularize compared to a DL model (which could have untold millions of possible parameter combinations, let alone getting the thing to train to a global optimum).
I think if you spent months getting your data and model structure to a good place, you could certainly get a DL model to out-perform a gradient boosted tree. But why do that, when the GBT will be done today?
- jb_s 4y agodo you reckon it's possible to somehow transfer learn from a GBT to a NN ?
- a-dub 4y agothis is along the lines of my thinking. people organize and summarize data before throwing it into spreadsheets, where deep learning models do their thing by generating new representations from raw data. in a sense, most data in spreadsheets is compressed and deep learning models prefer to find their own compression that best suits the task at hand. or in human terms: "these spreadsheets are garbage. i can't work with this. can you bring me the raw data please?" :)
- deleted 4y ago[deleted]
- username_exists 4y agooccam's razor
- drzoltar 4y agoI think another aspect is that most modern GBT models prefer the entire dataset to be in memory, thereby doing a full scan of the data for each iteration to calculate the optimal split point. That’s hard to compete with if your batch size is small in a NN model.
- anothernewdude 4y agothey also do subsampling of the data though.
- a-dub 4y agothat's an interesting idea. but at the end of the paper they do an analysis of the effect of different hyperparameters for the nets with their dataset and find that the batch size doesn't seem to matter much. (although they're trying size ranges like [256, 512, 1024] as opposed to turning batching off entirely)
- alexcnwy 4y agoThe issue isn’t batch size as a parameter but rather needing to load the entire dataset into memory
- a-dub 4y ago> thereby doing a full scan of the data for each iteration to calculate the optimal split point > (although they're trying size ranges like [256, 512, 1024] as opposed to turning batching off entirely) > The issue isn’t batch size as a parameter but rather needing to load the entire dataset into memory what's stored in memory is an implementation detail. the key idea is that the tree algorithms are choosing an optimal based on the entire dataset, where sgd is working on small randomly chosen batches. turning off batching means computing gradients on the entire dataset instead. although the typical bottleneck in gpu computing is moving data to and from the gpu's workarea (which is probably why you mention memory), there is nothing theoretical that says these computations could not be implemented in a streaming manner.
- thesz 4y agoOh, you touched my favorite topic of whole dataset training. Take a look at [1] and go straight to the page 8, figure 2(b). [1] http://proceedings.mlr.press/v48/taylor16.pdf http://proceedings.mlr.press/v48/taylor16.pdf The paper talks about whole dataset training and one of the datasets used is HIGGS [2]. The figure 2(b) shows two whole dataset training approaches (L-BFGS and ADMM) vs SGD. SGD tops at the accuracy with which both whole dataset approaches start, basically. [2] https://archive.ics.uci.edu/ml/datasets/HIGGS# https://archive.ics.uci.edu/ml/datasets/HIGGS# HIGGS is strange dataset. It is narrow, having only 29 features. It is also relatively long, about 11M samples (10M to train, 0.5M to validate and last 0.5M to test). It is also hard to get right with SGD. But if you perform whole dataset optimization, even linear regression can get you good accuracy [3] (some experiments of mine). [3] https://github.com/thesz/higgs-logistic-regression https://github.com/thesz/higgs-logistic-regression
- ellisv 4y agoI agree. The majority of DL layers are about feature engineering, not performing classification.
- wpietri 4y agoCould you say more about this? One of the things that interests me about nominal AI applications is the extent to which they're sort of a Mechanical Turk or what I've heard called Artificial Artificial Intelligence. By which I mean it's sold as computer magic, but most of the magic is actually humans sneaking in a lot of human judgement. That can come through humans directly massaging the output or through human-driven selection of results. But I've also been wondering to what extent natural human intelligence is getting put in at the lower layers of the system, like feature engineering.
- jb_s 4y agolooks like AI, quacks like a bunch of linear equations
- eru 4y agoLots of linear functions with \x -> max(0, x) thrown in between. (That's literally what neural nets with relu as activation unit do.)
- laichzeit0 4y agoThis is how I see it: Statistical modeling: input -> feature extraction (manual) -> model selection (manual) -> output Machine learning: input -> feature extraction (manual) -> model selection (auto) -> output Deep learning: input -> feature extraction (auto) -> model selection (auto) -> output So take a DL image classifier. The convolution + pooling layers perform automatic feature extraction. Back to OPs point, why use something like DL when you've already engineered your features?
- wpietri 4y agoWow, what a good summary. Thanks!
- oofbey 4y agoI think you’re on the right track that trees are good at feature engineering. But the key problem is that DL researchers are horrible at feature engineering, because they have never had to do it. These folks included. The feature engineering they do here is absolutely horrible! They use a QuantileTransform and that’s it. They don’t even tune the critical hyper parameter of the number of quantiles. Do they always use the scikitlearn default of 1,000 quantiles? No wonder uninformative features are hurting- they are getting expanded into 1000 even more uninformative features! Also with a single quantile transform like that, the relative values of the quantiles are completely lost! If the values 86 and 87 fall into different bins, the model has literally no information that the two bins are similar to each other, or even that they come from the same raw input. For a very large dataset a NN would learn its way around this kind of bone headed mistake. But for this size dataset, these researchers have absolutely crippled the nets with this thoughtless approach to feature engineering. There is plenty more to criticize about their experiments, but it’s probably less important. E.g. Their HP ranges are too small to allow for the kind of nets that are known to work best in the modern era (after Double Descent theory has been worked out) - large heavily regularized nets. They don’t let the nets get very big and they don’t let the regularization get nearly big enough. So. Bad comparison. But it’s also very true that XGB “just works” most of the time. NN’s are finicky and complicated and very few people really understand them well enough to apply them to novel situations. Those who do are working on fancy AI problems, not writing poor comparison papers like this one.
- eutectic 4y agoI think you misunderstand what quantiletransformer does; it just transforms the distribution of the data be be more normal, it's not a binning technique. To many quantiles will just result in excess noise, not loss of information.
- oofbey 4y agoOh I see. Interesting.
- btown 4y agoIt occurs to me that a system, trained on peer-reviewed applied-machine-learning literature and Kaggle winners, that generates candidates for structured feature-engineering specifications, based on plaintext descriptions of columns' real-world meaning, should be considered a requisite part of the "meta" here. Ah, and then you could iterate within the resulting feature-engineering-suggestion space as a hyper-parameter between experiments, which could be optimized with e.g. https://github.com/HIPS/Spearmint https://github.com/HIPS/Spearmint . The papers write themselves!
- lr1970 4y ago> I have a theory - tree based models require minimal feature engineering. Actually, the whole premise of Deep Learning is to learn proper feature representations from data with minimal data preprocessing. And it works wonderfully in CV and NLP but is less performant in tabular data. The paper indicates that there are several contributing factors to the DL underperforming.
- thom 4y agoWhat are the principled ways that tree based models handle categorical data? If you end up having to do one-hot encoding it feels like you need very wide forests or very deep trees. If your categorical data is actually vaguely continuous then splits can be quite efficient but that’s rare. I assume some day someone will be able to explain all this in information theoretic terms. I’m never sure if we’re comparing like with like (are the deep learning models we’re comparing against actually that deep, for example?) but clearly there’s something to the intuition that many small overfit models are more efficient than one big general model.
- coffee_am 4y agoOne way is to use categorical set splits [1] (proposed for categorical set inputs, but works for categorical features as well), used in TF-DF [1]. Greedy, and expensive to train (cheap inference though), but it gives great results. [1] https://arxiv.org/pdf/2009.09991.pdf https://arxiv.org/pdf/2009.09991.pdf [2] https://www.tensorflow.org/decision_forests/text_features https://www.tensorflow.org/decision_forests/text_features
- CapmCrackaWaka 4y agoSome implementations will group the categories into two groups, right and left. An easy way to do this is to create the groups based on whatever grouping creates the biggest decrease in the loss. There are more complicated implementations, lightgbm groups by the sum(gradient) / (sum(hessian) + smooth) where smooth is some smoothing parameter which prioritizes class size over loss decrease. They link to this paper in their docs: https://www.tandfonline.com/doi/abs/10.1080/01621459.1958.10501479 https://www.tandfonline.com/doi/abs/10.1080/01621459.1958.10...