4 ms·
It seems like learned positional encodings would still prevent you from doing fine tuning on a larger context size, though, so maybe using alibi is still releva
by oneseven 3y ago
It seems like learned positional encodings would still prevent you from doing fine tuning on a larger context size, though, so maybe using alibi is still relevant (although I have not read that paper).
- jimsimmons 3y agoYou can collapse all positions beyond a length to a specific bucket like T5