7 ms·
I recently had the same questions and here is how I understand it: 1. You could concatenate the positional embedding and the semantic embedding and that way is
by samvher 3y ago
I recently had the same questions and here is how I understand it:
1. You could concatenate the positional embedding and the semantic embedding and that way isolate them from each other. But if that separation is necessary, the model can learn the separation itself as well (it can make positional embeddings and semantic embeddings orthogonal to each other), so using addition is strictly more general.
2. My sense is that you could merge the Q and K matrices and everything would work mostly the same, but with multi-headed attention this will typically result in a larger matrix than the combined sizes of Q and K. It's basically a more efficient matrix factorization.
Curious to see if I got this right and if there is more to it.
- tlb 3y agoYes, that’s my understanding. One advantage of summing is that the lower frequency terms hardly change for a small text, so effectively there is more capacity for embeddings with short texts, while still encoding order in long texts.