3 ms·
Memory does not scale quadratically with sequence length.
by why_only_15 3y ago
Memory does not scale quadratically with sequence length.
- RC_ITR 3y agoDuring training, you have to store a dot product of Q and V that has dimension Ncrt^2. That's quadratic scaling, no?
- why_only_15 3y agopresumably you mean a dot product of Q and K, and no you do not have to store this: https://arxiv.org/abs/2205.14135 https://arxiv.org/abs/2205.14135
- RC_ITR 3y agoI mean, sure you can work around it, but from your own link: >since the time and memory complexity of self-attention are quadratic in sequence length
- why_only_15 3y agoExcept in practice this is not true, and hasn't been for more than a year. It's not just a workaround either -- FlashAttention is both faster at runtime and uses less memory.
- deleted 3y ago[deleted]