5 ms·
250k because thats what the smaller models supported and theres better training data in that part of the window there have been some papers suggesting that the
by 8note 2mo ago
250k because thats what the smaller models supported and theres better training data in that part of the window
there have been some papers suggesting that the useful context is even smaller, and stays fixed as you change the context window size.
as a more general case though, i think the possibilities for what youd need to include in training to have many different paths of text be well represented enough over the long window means the later tokens will almost always be a lot more random than early ones?
there's a sheer amount of bits problem.
- charcircuit 2mo ago>in that part of the window What matters is the relative distance. Using the latter 250k should be equivalent to using the first 250k.