2 ms·
Thank you! Great question. "Infinite-length sequence processing" in StreamingLLM refers to handling much longer sequences than the model's training window (e.g
by t55 2y ago
Thank you! Great question.
"Infinite-length sequence processing" in StreamingLLM refers to handling much longer sequences than the model's training window (e.g., millions of tokens), by combining a sliding window for recent tokens with fixed attention sinks from the start of the sequence.
I can't speak for DeepSeek, but if I had to guess, I'd say that the infinite context window isn’t practical because storing all past tokens eventually becomes too expensive.