3 ms·
You mean "Multiply the vectors by the other vectors. This is attention - it's the magic of transformers, that enables combining information from multiple token
by vikp 3y ago
You mean "Multiply the vectors by the other vectors. This is attention - it's the magic of transformers, that enables combining information from multiple tokens together. This generates a new matrix."?
It's really oversimplified, as I mentioned. A more granular look is:
- Project the vectors with a linear regression. In decoder-only attention (what we usually use), we project the same vectors twice with different coefficients. We call the first projection queries, and the second keys. This transforms the vectors linearly.
- Find the dot product of each query vector against the key vectors (multiply them)
- (training only) Mask out future vectors, so a token can't look at tokens that come after it
- At this point, you will have a matrix indicating how important each query vector considers each other vector (how important each token considers the other tokens)
- Take the softmax, which both ensures all of the attention values for a vector sum to 1, and penalizes small attention values
- Use the softmax values to get a weighted sum of tokens according to the attention calc.
- This will turn one vector into the weighted sum of the other vectors it considers important.
The goal of this is to incorporate information from multiple tokens into a single representation.