4 ms·
Landmark Attention: Random-Access Infinite Context Length for Transformers
- version_five 3y agoThey say it's infinite but then they use it to increase Llama's context length to 32k tokens instead of infinity. Still seems interesting.
- two_in_one 3y agoActually cool if it works. This would make many things possible on a single desktop. I'm afraid landmarking is actually a lossy compression. And the real content's size still limits model's even theoretical abilities.
- MacsHeadroom 3y agoIt's infinite but not free. Larger context still means more VRAM used and longer compute times.
- viraptor 3y agoThe link to the repo (https://github.com/epfml/landmark-attention https://github.com/epfml/landmark-attention) leads to "we'll publish something later".