3 ms·
Oh interesting, didn't know. How does this work past the first transformer in the stack?
by Scene_Cast2 2y ago
Oh interesting, didn't know. How does this work past the first transformer in the stack?
- lonk11 2y agoMy understanding is that the attention in all transformer layers is "causal" - that is the output of a transformer layer for token N depends only on tokens from 0 to N. This means that every attention layer can use previously calculated outputs for the same prompt prefix. So it only needs to calculate from scratch starting from the first unique token in the prompt sequence.
- danielmarkbruce 2y agoI had the same question... my guess is you can do a layer by layer cache. Ie a cache in the first layer, then another independent second layer cache, and so on.