4 ms·
The method they use is surprisingly simple. They claim GPTs can’t effectively generate beyond the context window because our models overfit on positional encodi
by valine 3y ago
The method they use is surprisingly simple. They claim GPTs can’t effectively generate beyond the context window because our models overfit on positional encodings. The fix is literally to cap the positional encodings at inference time.
It makes sense intuitively that the exact position of tokens really only matters for adjacent or near adjacent tokens. For far away tokens a rough position is fine.
- nulld3v 3y agoRoPE scaling has been a thing for a while already: https://arxiv.org/abs/2306.15595 https://arxiv.org/abs/2306.15595 Does anybody know what the difference is between the approach in OP vs other RoPE scaling approaches?
- hexaga 3y agoThe core problem is: there's not enough unique, trained positions. Naively going past the end of training ctx makes you run straight into out of distribution positions, and things become incoherent. For a model trained with a ctx size of 2, that looks like: `[0, 1, *incoherence starts* 2, 3]` Existing RoPE scaling methods try to stay in-distribution by assigning positions between the known in-distribution ones: `[0, 0.5, 1, 1.5]`. This is still ~kinda OOD, but works w/ some fine tuning. The method in OP breaks the core premise that we need unique positions at all, and just gives multiple tokens the same position: `[0, 0, 1, 1]`.
- cma 3y agoWouldn't that mean (until higher level embeddings) compound phrases far away are unordered? And numbers are fragmented by token boundary and scrambled up?
- valine 3y agoEmbeddings at lower layers aren’t going to be looking very far beyond nearby or adjacent embeddings as they refine their meaning. For a number like 3.14, the tokens 3 and 14 are important to each other, but entirely unimportant to understand the meaning of a question later in the context. It’s only at later layers that an embedding representing the concept of PI becomes important to the question embeddings. As I understand it the positional encodings are calculated relative to the token in question. It’s not like 3 and 14 are unordered tokens from their own perspective.