5 ms·
We introduce StreamingLLM, an efficient framework that enables LLMs trained with a finite length attention window to generalize to infinite sequence length with
by guywithabowtie 3y ago
We introduce StreamingLLM, an efficient framework that enables LLMs trained with a finite length attention window to generalize to infinite sequence length without any fine-tuning. We show that StreamingLLM can enable Llama-2, MPT, Falcon, and Pythia to perform stable and efficient language modeling with up to 4 million tokens and more.
- stavros 3y agoSorry, what does "up to 4 million tokens and more" mean? It seems like a contradiction.
- jamesblonde 3y agoHere's a reference describing what a context window for LLMs is: https://www.hopsworks.ai/dictionary/context-window-for-llms https://www.hopsworks.ai/dictionary/context-window-for-llms
- catskul2 3y agoNot really a contradiction so much as redundant/poorly worded. Should have said, "at least 4 million tokens".
- deleted 3y ago[deleted]