3 ms·
I think for O(N^2) transformer inference you need to cache all the activations.
by eutectic 3y ago
I think for O(N^2) transformer inference you need to cache all the activations.
- thomasahle 3y agoYou only need to cache the key/value pairs. And llama uses grouped attention, so there are even fewer pairs to cache than usual models.