4 ms·
I'm no expert, but I think it's so that the model can learn the relative position wrt other tokens. They use indices for models like vision transformers with a
by tipsytoad 3y ago
I'm no expert, but I think it's so that the model can learn the relative position wrt other tokens.
They use indices for models like vision transformers with a fixed number of patches but for variable length context I think it's more beneficial to use encodings that can also capture the relative distance.