5 ms·
DL practitioner for a decade here: The OP doesn’t explain anything. It just vaguely talks about a few things that might break when scaling context. But that me
by jimsimmons 3y ago
DL practitioner for a decade here:
The OP doesn’t explain anything. It just vaguely talks about a few things that might break when scaling context. But that means nothing.
Take for example sinusoidal embeddings they talk about. Of course it breaks for large contexts but in no one uses it. GPT uses learned positional embeddings so the entire section is irrelevant.
Copy this for pretty much everything else.
Being an expert in a field has never been this exhausting
- it_citizen 3y agoTry virologist 3 years ago.
- oneseven 3y agoIt seems like learned positional encodings would still prevent you from doing fine tuning on a larger context size, though, so maybe using alibi is still relevant (although I have not read that paper).
- jimsimmons 3y agoYou can collapse all positions beyond a length to a specific bucket like T5