4 ms·
I'm pretty sure in GPT3.5+ models, this concept of attention holds no longer true. In [1] the author suggests they are using "Intra-tensor sparsity" from a 2021
by spi 3y ago
I'm pretty sure in GPT3.5+ models, this concept of attention holds no longer true. In [1] the author suggests they are using "Intra-tensor sparsity" from a 2021 paper from Google [2].
Details aside, math suggests they _must_ be using some sparse attention method: the memory used by attention is O(l^2 * d) (here O notation is a bit of overkill: it's exactly that number in 8 bit quantization or twice that in 16 bit floats). With l=32k, and d probably in the order of a few k (even "old" BERT model had it at 768, it only went up since), that would be of the order of a few TB. For (a piece of) _a single_ layer, they probably have dozens, if not hundreds, of it. The largest GPUs in commerce have 80GB of memory; there's no way they are really using that many GPUs for each single layer (even leaving aside the fact that runtime would probably be absolutely horrible).
Implicitly or explicitly, the transformer must learn to have both a "summarized" version of attention (that tells it where to focus on) and more "detailed" ones. It doesn't need to be specifically coded, if using sparse attention, the model might for example get different levels of attentions at different layers, somehow, but they probably have coded something a bit smarter than that.
[1] https://kir-gadjello.github.io/posts/gpt4-some-technical-hypotheses/ https://kir-gadjello.github.io/posts/gpt4-some-technical-hyp...
[2] https://arxiv.org/abs/2111.12763 https://arxiv.org/abs/2111.12763