4 ms·
> For a sequence with N = 100,000 tokens, it would mean cost dropping by a factor of 100,000× I'm not sure I understand the intended interpretation of this. Co
by chillee 3y ago
> For a sequence with N = 100,000 tokens, it would mean cost dropping by a factor of 100,000×
I'm not sure I understand the intended interpretation of this. Concretely speaking, if it cost CoolAI 100k seconds of compute to process a sequence of length 100k, it would not take them 1 second now.
I agree that as sequences become longer, the quadratic component will become more important. But as models get bigger, the attention component also becomes less important.
For example, to take a concrete model (say Llama-70B), it takes about 1.4e16 MLP FLOPs (70 billion * 100000 * 2) to process 100k tokens. The attention component takes about 6.5e15 FLOPS (80 [layers] * 100k [sequence length]^2 * 8192 [hidden dim]).
So even if attention turned constant it would reduce runtime by about 30% with today's model at 100k sequence length.
- cs702 3y ago> I'm not sure I understand the intended interpretation of this. Concretely speaking, if it cost CoolAI 100k seconds of compute to process a sequence of length 100k, it would not take them 1 second now. Cost would be lower by a factor of sequence length N = 100,000. If you call the factor C, cost would be lower by C × N. For most models and context lengths today, C is below 1. As we continue to increase sequence length N -- say, as we go from 100K to 100M tokens, incorporating multiple modalities -- factor C will increase toward 1 for all Transformers. Again, I see how what I wrote could be misinterpreted. Sorry about that!