5 ms·
There's definitely xor type nonlinearities going on in factor models of personality types. The interaction between two factors is often very different at the ex
by seertaak 3y ago
There's definitely xor type nonlinearities going on in factor models of personality types. The interaction between two factors is often very different at the extrema of each factor - often flipping sign.
Regression models can't capture that. Now it's possible that the behavioral nonlinearities are somehow emergent, and that despite those, the ECG can be captured by a linear model. But I wouldn't want to bet on that.
- Calavar 3y agoNonlinearities in ML existed long before DL and ReLU. What's at the end of a deep CNN? Probably global pooling followed by a dense layer? That final dense layer is taking a weighted combination of the last stack of feature maps. So deep learning is fundamentally the same as the old approach: Construct features with a nonlinear transformation of the input. Then calculate a score by taking a linear combination of those features. The only difference with deep learning is that the manner of constructing those nonlinear features is left unspecified.
- seertaak 3y agoInteresting comment. I'm not an expert in CNNs but what you're saying makes a lot of sense. Question: are you saying the final layer is "in effect" a linear combination? At least in transformer architectures, the dense end block is iirc three layers deep and uses relu. Even if the CNN's dense final part is one layer deep, wouldn't it also be using relu activations? Even a one layer dense ANN can capture nonlinearities if it has a nonlinear activation fn. But maybe in practice the activations don't do a lot of work? Or am I simply mistaken about final layer, does it simply have linear activations? Also, can you share your intuition about the nonlinear features? Spectral/wavelet analysis? Or something more complex?
- Calavar 3y ago> Question: are you saying the final layer is "in effect" a linear combination? Not just in effect a linear combination; it is a linear combination. There are some exotic nonlinear NN layers that are used in particular niches. But in general, NN layers are syntax sugar over matrix multiplications (i.e. linear functions). The nonlinearities are the activation functions between the layers. > But maybe in practice the activations don't do a lot of work? No, the activations do do a lot of work. There are real nonlinearities in the network. > Or am I simply mistaken about final layer, does it simply have linear activations? Sometimes you actually do use a linear activation on the final layer for certain regression tasks, but that's not the main thing I'm getting at. Let's say that your final layer has is a ReLU activation. What is this conceptually? You are taking a linear combination of the features from the previous layer and them clamping the result to >= 0. Sure, that's a nonlinearity, but it isn't going to have much in the way of emergent modeling capabilities. You need to stack many, many nonlinearities before you get that. So my point is that a deep neural net of N layers boils down to a complex nonlinear function of N - 1 layers, followed by a "dumb" linear combination in the final layer. You can do this with traditional ML methods as well, but you have to handcraft your nonlinearities. > Also, can you share your intuition about the nonlinear features? Spectral/wavelet analysis? Or something more complex? There's innumerable possibilities here. It could be starting with a method that's inherently nonlinear, like nonlinear PCA, polynomial regression. Or it can involve transforming the output of a linear function (like the Fourier transform) in a nonlinear way. Admittedly this very tough. And for really tough problems, like video synthesis, effectively impossible. But NNs get thrown at much simpler problems all the time.