3 ms·
good news - we've actually included optimized causal and non-causal versions of the flash attention backwards pass with TK - would love for you to check them ou
by Aaryan44 2y ago
good news - we've actually included optimized causal and non-causal versions of the flash attention backwards pass with TK - would love for you to check them out!
causal: https://github.com/HazyResearch/ThunderKittens/blob/main/examples/attn_causal/h100_train.cu https://github.com/HazyResearch/ThunderKittens/blob/main/exa...
non-causal: https://github.com/HazyResearch/ThunderKittens/blob/main/examples/attn/h100/h100_train.cu https://github.com/HazyResearch/ThunderKittens/blob/main/exa...
- lucidrains 2y agoamazing work! thank you!
- Aaryan44 2y agoThanks @lucidrains :)
- pama 2y agoAwesome. Do you happen to have a benchmark against the latest (v9.1) cuDNN implementation?
- Aaryan44 2y ago@pama, if useful - here are utilization numbers for our attention backwards kernels (causal and non-causal, head dim = 64): https://github.com/HazyResearch/ThunderKittens/blob/main/attn.png https://github.com/HazyResearch/ThunderKittens/blob/main/att...