3 ms·
The core problem is: there's not enough unique, trained positions. Naively going past the end of training ctx makes you run straight into out of distribution po
by hexaga 3y ago
The core problem is: there's not enough unique, trained positions. Naively going past the end of training ctx makes you run straight into out of distribution positions, and things become incoherent.
For a model trained with a ctx size of 2, that looks like: `[0, 1, *incoherence starts* 2, 3]`
Existing RoPE scaling methods try to stay in-distribution by assigning positions between the known in-distribution ones: `[0, 0.5, 1, 1.5]`. This is still ~kinda OOD, but works w/ some fine tuning.
The method in OP breaks the core premise that we need unique positions at all, and just gives multiple tokens the same position: `[0, 0, 1, 1]`.