3 ms·
I'm very much not an expert either but apparently for the regular "attention" mechanism memory and compute requirements scale quadratically with respect to inpu
by ldhough 4y ago
I'm very much not an expert either but apparently for the regular "attention" mechanism memory and compute requirements scale quadratically with respect to input sequence length. So increasing it by just two orders of magnitude would mean (I think) a full context window needs 10000x more memory and compute time and presumably costs would go up by at least as much. GPT3 (I think) uses the regular attention mechanism while 4 is unknown.
However, GPT4 claims there are techniques to improve scaling (complexity down to sub-quadratic or linear) without affecting accuracy too much (I have no clue if true): sparse attention, long-range arena, reformer, and performer.
I'm also pretty sure I've read (and anecdotally it seems true) that accuracy decreases with longer input/output sequences regardless. How much I also don't know.