5 ms·
Not the parent, but NNs typically work better when you can't linearize your data. For classification, that means a space in which hyperplanes separate classes,
by queuebert 3y ago
Not the parent, but NNs typically work better when you can't linearize your data. For classification, that means a space in which hyperplanes separate classes, and for regression a space in which a linear approximation is good.
For example, take the circle dataset here: https://playground.tensorflow.org https://playground.tensorflow.org
That doesn't look immediately linearly separable, but since it is 2D we have the insight that parameterizing by radius would do the trick. Now try doing that in 1000 dimensions. Sometimes you can, sometimes you can't or don't want to bother.
- mjhay 3y agoThat's an advantage over linear models, but GBTs handle non linearly-separated data just fine. Each individual tree can represent an arbitrary piecewise-constant function given enough depth, and then each tree in turn tries to minimize the loss on the residual of the previous trees. As such, they're effectively like a neural network with two hidden layers in terms of expressiveness.
- melondonkey 3y agoThis explanation doesn’t make sense to me. What do you mean by “linearize your data”—tree methods assume no linear form and are not even monotonically constrained. Classification is not done by plane-drawing but by probability estimation + cost function
- dist-epoch 3y agoA tree split can be considered plane-drawing.
- CuriouslyC 3y agoNote that if linear separability is the only issue you can just use kernel methods. In fact, gaussian processes are equivalent to a single hidden layer neural network with infinite hidden values. The magic of deep neural networks comes from modeling complicated conditional probability distributions, which lets you do generative magic but isn't going to give you significantly better results than ensemble kNN when you're discriminating and the conditional distribution is low variance. Ensemble methods are like a form of regularization and they also act as a weak bootstrap to better model population variance, so it's no surprise that when they're capable of modeling the domain, they perform better than unregularized, un-bootstrapped neural network model. There are still tons of situations where ensemble methods can't model the domain, and if you incorporated regularization and bootstrapping into a discriminative NN model it would probably perform equivalently to the ensemble model.