4 ms·
As I understand it, the main result of the paper relies on these 4 assumptions: - That the dimensionality of the output of the network is smaller than that of
by avallet 10y ago
As I understand it, the main result of the paper relies on these 4 assumptions:
- That the dimensionality of the output of the network is smaller than that of the input. That is usually the case in image recognition, where the image is width x height x channels dimensional, while the output is usually a much smaller number of label-wise probabilities. It probably isn't the case when you generate data from some smaller representation, e.g. with autoencoders, image generation, etc.
- That the input data is decorrelated, and that the input data is uncorrelated with the output ground truth. The former can easily be obtained via a whitening transformation in many cases in practice. I am not quite sure about the latter.
- That whether a connection in the network is activated or not is random with the same probability of success across the network. Active means the ReLU activation function has output greater than 0. Many people initialize weights in the network with some 0-mean random variable and some constant bias, in which case that assumption should hold true at the beginning of training. Whether that assumption holds throughout training could easily be verified empirically - i.e. by monitoring the network's activation.
- That the network activations are independent of the input, the weights and each other. That's obviously not completely true - the network activations of a given layer are a function of the previous layer's activations and weights, and ultimately of the input in the first layer. With large enough networks, this may hold sufficiently in practice - any single activation should not depend very significantly on any other single variable.
I may have missed something in interpreting the maths, any comment is appreciated. From a practical standpoint, especially for computer vision, these assumptions seem quite reasonable. I am not qualified however to comment on the proof of this result, so I would wait on peer review. Still, it is heartening to see the theory of deep learning finally catching up with practice!
- imh 10y ago>..and that the input data is uncorrelated with the output ground truth. I haven't read it yet. Is that just linear correlation, or full independence? If it's independence, then there's no signal, right?
- avallet 10y agoLinear correlation. Actually, it seems to be a bit more general than just uncorrelated, i.e. if the input is a m by n matrix X and ground truth a k by n matrix Y, the author requires that XX^T and XY^T to be full-rank. A whitening transformation would yield the identity matrix for XX^T, but that's a bit stronger than what's strictly necessary. My interpretation of XY^T being full-rank meaning X and Y being uncorrelated might indeed be mistaken.
- pedrosorio 10y agoCould you clarify what is the mathematical definition of "X and Y are uncorrelated" for two matrices?
- avallet 10y agoI am not quite sure there is such a thing. :p I was playing a bit loose with the mathematics here, and trying to find some more intuitive way to explain "XY^T is full-rank", but it got confusing. Sorry about that. I will edit my initial post accordingly. (Ah, can't edit it seems, oh well)
- conjectures 10y agoIt is mistaken, if X=Y=I then XY' is full rank but each column in X is a linear function of its counterpart in Y.