3 ms·
I was under the impression that each new token attends to one previous token per attention head, and that the slowdown observed was more because those attended
by TomatoCo 23d ago
I was under the impression that each new token attends to one previous token per attention head, and that the slowdown observed was more because those attended to tokens are more spread out in memory and get less memory-architecture-style cache hits.