4 ms·
The fundamental operation done by the transformer, softmax(Q.K^T).V, is essentially a KV-store lookup. The Query is dotted with the Key, then you take the soft
by jacobn 3y ago
The fundamental operation done by the transformer, softmax(Q.K^T).V, is essentially a KV-store lookup.
The Query is dotted with the Key, then you take the softmax to pick mostly one winning Key (the Key closest to the Query basically), and then use the corresponding Value.
That is really, really close to a KV lookup, except it's a little soft (i.e. can hit multiple Keys), and it can be optimized using gradient descent style methods to find the suitable QKV mappings.
- naveen99 3y agoNot sure there is any real lookup happening. Q,K are the same and sometimes even v is the same…
- toxik 3y agoQ, K, V are not the same. In self-attention, they are all computed by separate linear transformation of the same input (ie the previous layer’s output). In cross-attention even this is not true, then K and V are computed by linear transformation of whatever is cross-attended, and Q is computed by linear transformation of the input as before.