3 ms·
would be interested to see thunderkittens (great name!) tackle the flash attention backwards pass, which is an order of magnitude harder than the forward
by lucidrains 2y ago
would be interested to see thunderkittens (great name!) tackle the flash attention backwards pass, which is an order of magnitude harder than the forward
- Aaryan44 2y agogood news - we've actually included optimized causal and non-causal versions of the flash attention backwards pass with TK - would love for you to check them out! causal: https://github.com/HazyResearch/ThunderKittens/blob/main/examples/attn_causal/h100_train.cu https://github.com/HazyResearch/ThunderKittens/blob/main/exa... non-causal: https://github.com/HazyResearch/ThunderKittens/blob/main/examples/attn/h100/h100_train.cu https://github.com/HazyResearch/ThunderKittens/blob/main/exa...
- lucidrains 2y agoamazing work! thank you!
- Aaryan44 2y agoThanks @lucidrains :)
- pama 2y agoAwesome. Do you happen to have a benchmark against the latest (v9.1) cuDNN implementation?
- Aaryan44 2y ago@pama, if useful - here are utilization numbers for our attention backwards kernels (causal and non-causal, head dim = 64): https://github.com/HazyResearch/ThunderKittens/blob/main/attn.png https://github.com/HazyResearch/ThunderKittens/blob/main/att...