4 ms·
> Are these deep learning detectors better than running a simple regression after using something like a wavelet transform to detect these specific features? I
by 4death4 3y ago
> Are these deep learning detectors better than running a simple regression after using something like a wavelet transform to detect these specific features?
It’s pretty insane to think that a decent neural net trained on sufficient data wouldn’t be able to outperform basic regression.
- ssivark 3y agoWhat makes you think that a neural net must necessarily be significantly better? How much better is significant? Can one actually train a large enough network on (finite) available data? Sometimes it might take god awful amounts of data+compute+layers to rediscover transforms that we have ready access to, once math has matured enough.
- 4death4 3y agoIt’s logically obvious. You can recreate any linear regression using a neural network. So a neural network approach will always be at least as good as linear regression. But a neural network can model non-linear relationships as well. Now ask yourself, what is the likelihood of there being a non-linear relationship in the data? As with most real-world data, the likelihood of non-linear relationships is extremely high. That’s why we can be quite certain a neural network approach will able to outperform linear regression.
- Calavar 3y ago> It’s logically obvious. You can recreate any linear regression using a neural network. So a neural network approach will always be at least as good as linear regression. Well, maybe it's logically obvious, but it's also empirically false: https://www.youtube.com/watch?v=x7psGHgatGM https://www.youtube.com/watch?v=x7psGHgatGM (Skip to 13:36 to get to the point) > But a neural network can model non-linear relationships as well. Now ask yourself, what is the likelihood of there being a non-linear relationship in the data? As with most real-world data, the likelihood of non-linear relationships is extremely high. That’s why we can be quite certain a neural network approach will able to outperform linear regression. As I've said elsewhere in this thread, a deep neural network of N layers can be decomposed into a complex nonlinear function of N - 1 layers followed by a linear combination of the nonlinear features generated that function. There's no law that says that you have to use a neural network to generate your nonlinear features. You can use any method you like and then linearly combine those.
- 4death4 3y agoI don’t think you understand what you’re saying. “A linear combination of non-linear features” is not what people mean when they talk about linear regression. And even if they were, a neural network will do a much better job of generating the non-linear features than you will be hand. So again, it’s logically obvious that a neural network will be capable of performing better than simple linear regression.
- Calavar 3y ago> I don’t think you understand what you’re saying. “A linear combination of non-linear features” is not what people mean when they talk about linear regression. According to who? Did you watch the talk I linked? Because in that video, NeurIPS gave their test of time award to a paper where they generated nonlinear features and used those to train a linear classifier. I guess the committee at NeurIPS also didn't understand what they are saying? > And even if they were, a neural network will do a much better job of generating the non-linear features than you will be hand. So again, it’s logically obvious that a neural network will be capable of performing better than simple linear regression. Again, logically obvious but often empirically false.
- Calavar 3y agoIt's really not. If you throw a blob of unstructured data at a neural network, then sure, it will beat everything else by a mile. If you actually try to understand your data and engineer your features, you can often beat DL with much simpler methods. (Obviously YMMV depending on the specific problem domain.) Unfortunately, feature detection and feature selection are dying arts. Just throw all the data at a CNN or transformer, report the accuracy, and that's it, that's the paper. No analysis of robustness, edge cases, etc. because you can't easily do those things if you haven't taken the time to understand your data in the first place.
- seertaak 3y agoThere's definitely xor type nonlinearities going on in factor models of personality types. The interaction between two factors is often very different at the extrema of each factor - often flipping sign. Regression models can't capture that. Now it's possible that the behavioral nonlinearities are somehow emergent, and that despite those, the ECG can be captured by a linear model. But I wouldn't want to bet on that.
- Calavar 3y agoNonlinearities in ML existed long before DL and ReLU. What's at the end of a deep CNN? Probably global pooling followed by a dense layer? That final dense layer is taking a weighted combination of the last stack of feature maps. So deep learning is fundamentally the same as the old approach: Construct features with a nonlinear transformation of the input. Then calculate a score by taking a linear combination of those features. The only difference with deep learning is that the manner of constructing those nonlinear features is left unspecified.
- seertaak 3y agoInteresting comment. I'm not an expert in CNNs but what you're saying makes a lot of sense. Question: are you saying the final layer is "in effect" a linear combination? At least in transformer architectures, the dense end block is iirc three layers deep and uses relu. Even if the CNN's dense final part is one layer deep, wouldn't it also be using relu activations? Even a one layer dense ANN can capture nonlinearities if it has a nonlinear activation fn. But maybe in practice the activations don't do a lot of work? Or am I simply mistaken about final layer, does it simply have linear activations? Also, can you share your intuition about the nonlinear features? Spectral/wavelet analysis? Or something more complex?