4 ms·
RoPE scaling has been a thing for a while already: https://arxiv.org/abs/2306.15595 https://arxiv.org/abs/2306.15595 Does anybody know what the difference is b
by nulld3v 3y ago
RoPE scaling has been a thing for a while already: https://arxiv.org/abs/2306.15595 https://arxiv.org/abs/2306.15595
Does anybody know what the difference is between the approach in OP vs other RoPE scaling approaches?
- hexaga 3y agoThe core problem is: there's not enough unique, trained positions. Naively going past the end of training ctx makes you run straight into out of distribution positions, and things become incoherent. For a model trained with a ctx size of 2, that looks like: `[0, 1, *incoherence starts* 2, 3]` Existing RoPE scaling methods try to stay in-distribution by assigning positions between the known in-distribution ones: `[0, 0.5, 1, 1.5]`. This is still ~kinda OOD, but works w/ some fine tuning. The method in OP breaks the core premise that we need unique positions at all, and just gives multiple tokens the same position: `[0, 0, 1, 1]`.