4 ms·
For those who don't know the term "kernel smoothing", it just means ∑ᵢ yᵢ · K(xᵢ, xₒ) ⁄ (∑ⱼ K(xⱼ, xₒ)) In regular attention, we let K(xᵢ, xₒ) = exp(<xᵢ, x
by thomasahle 1y ago
For those who don't know the term "kernel smoothing", it just means
∑ᵢ yᵢ · K(xᵢ, xₒ) ⁄ (∑ⱼ K(xⱼ, xₒ))
In regular attention, we let K(xᵢ, xₒ) = exp(<xᵢ, xₒ>).
Note that in Attention we use K(qᵢ, kₒ) where the q (query) and k (key) vectors are not the same.
Unless you define K(xᵢ, xₒ) = exp(<W_q xᵢ, W_k xₒ>) as you do in self-attention.
There are also some attention mechanisms that don't use the normalization term, (∑ⱼ K(xⱼ, xₒ)), but most do.
- throwup238 1y ago> ∑ᵢ yᵢ · K(xᵢ, xₒ) ⁄ (∑ⱼ K(xⱼ, xₒ)) That clarifies things...
- mrks_hy 1y agoIt does, for anybody who studied math at the level you need to understand Attention (some linear algebra). Please no low effort comments, ask if you don't have the math background and people will gladly help. This is sum notation, see https://en.m.wikipedia.org/wiki/Summation https://en.m.wikipedia.org/wiki/Summation
- auggierose 1y agoNo, it doesn't really clarify things. I had the best linear algebra grades in my year at my university, and if you don't know anything about kernels, this is not helpful (what are xi and yi in the first place?).
- lambdasquirrel 1y agoIt's all described in the referenced link. No need for everyone to get antsy. > 0. http://bactra.org/notebooks/nn-attention-and-transformers.ht http://bactra.org/notebooks/nn-attention-and-transformers.ht...