3 ms·
That's true in the realm of LLMs. But even in this case, the position information is added only into the first layer. Tokens in later layers can choose to "forg
by markisus 2y ago
That's true in the realm of LLMs. But even in this case, the position information is added only into the first layer. Tokens in later layers can choose to "forget" this information. In addition there are applications of transformers in other domains. See https://github.com/cvg/LightGlue https://github.com/cvg/LightGlue or https://facebookresearch.github.io/3detr/ https://facebookresearch.github.io/3detr/
- topwalktown 2y agoTransformers like Llama use rotary embeddings which are applied in every single attention layer https://github.com/huggingface/transformers/blob/222505c7e4d08da9095d12ddb72fb653f4b6da33/src/transformers/models/llama/modeling_llama.py#L275 https://github.com/huggingface/transformers/blob/222505c7e4d...
- markisus 2y agoVery interesting! Do you know if there were any studies about whether this improves performance?