3 ms·
I’m working on an inference platform that allows for tokens to be appended to the context after some tokens have been generated. If there’s other sequences in t
by zackangelo 2y ago
I’m working on an inference platform that allows for tokens to be appended to the context after some tokens have been generated. If there’s other sequences in the batch, it means they’ll have to be padded. Currently this means I can’t use FlashAttention because it doesn’t support arbitrary masks/padding masks… can ThunderKittens help me?