3 ms·
One thing you may have overlooked - table 5, the proof-pile table, only goes up to a 262k evaluation window (meaning - although the model has an extended contex
by Fripplebubby 3y ago
One thing you may have overlooked - table 5, the proof-pile table, only goes up to a 262k evaluation window (meaning - although the model has an extended context window of 2048k according to the method proposed, they are not feeding in that many tokens, only 262k tokens - so, about 13% of the total possible window).
Why? I think this is because books3 contains, you know, books - including some really long books, and proof-pile contains math papers and math stuff, which isn't as long.
So overall I think what you're seeing is a general trend of increasing perplexity on windows above 256k, between 256k-2048k, which is probably not so surprising - or at least, not so surprising when you consider the context of the paper, which is taking a model pre-trained with a much shorter context window and extending the context window using a novel technique. It's hard to adapt a model trained to do one thing into doing another thing, and that's what they're doing, so in that context, it tracks that the longer the context window, the worse the performance.