3 ms·
> Another think I simply did not get from the original Transformer paper is that the learning of self-attention happens in the linear layers. You can replace t
by PartiallyTyped 3y ago
> Another think I simply did not get from the original Transformer paper is that the learning of self-attention happens in the linear layers.
You can replace the KVQ kernels with any parametric computation that allows you to pass gradients through, and you will have learning.
I think some newer language model architectures use residual blocks here, and some vision transformers use FeedForward networks for the kernels.