4 ms·
Very clean writeup. On the attention sinks, you mention they enable "infinite-length sequence processing". What does that mean exactly in practice? Isn't deepse
by deepdarkforest 2y ago
Very clean writeup.
On the attention sinks, you mention they enable "infinite-length sequence processing". What does that mean exactly in practice? Isn't deepseek still capped at 128k?
- m348e912 2y agoAgreed on the writeup itself. It's beautifully written and presented. Kudos to Jean Kaddour and anyone else that may have been involved in putting it together.
- t55 2y agoThank you so much, glad you liked it
- t55 2y agoThank you! Great question. "Infinite-length sequence processing" in StreamingLLM refers to handling much longer sequences than the model's training window (e.g., millions of tokens), by combining a sliding window for recent tokens with fixed attention sinks from the start of the sequence. I can't speak for DeepSeek, but if I had to guess, I'd say that the infinite context window isn’t practical because storing all past tokens eventually becomes too expensive.