4 ms·
This was my first thought too. AFAIK each layer encodes different information, and it's not clear that the last layer would be able to communicate well with the
by mbowcut2 2y ago
This was my first thought too. AFAIK each layer encodes different information, and it's not clear that the last layer would be able to communicate well with the first layer without substantial retraining.
Like in a CNN for instance, if you fed later representations back in to the first kernels they wouldn't be able to find anything meaningful because it's not the image anymore, it's some latent representation of the image that the early kernels aren't trained on.