3 ms·
This is pretty great. PyTorch uses triton as the backend for torch.compile (the big feature of PyTorch 2.0, and the necessary part for making Flex Attention in
by Scene_Cast2 2y ago
This is pretty great. PyTorch uses triton as the backend for torch.compile (the big feature of PyTorch 2.0, and the necessary part for making Flex Attention in the about to be released 2.5 usably fast).
Triton's team doesn't support Windows, and, worse yet, does not accept community PRs to enable any sort of support.
Here's the github issue: https://github.com/triton-lang/triton/issues/1640 https://github.com/triton-lang/triton/issues/1640
And here's the performance comparison of Flex Attention with and without torch.compile (tldr it's 3x slower than a standard MHA when not compiled): https://github.com/rasbt/LLMs-from-scratch/blob/76e9a9ec02a1a060aac61608598fdd50cc7d52bd/ch03/02_bonus_efficient-multihead-attention/mha-implementations.ipynb https://github.com/rasbt/LLMs-from-scratch/blob/76e9a9ec02a1...
EDIT: after taking a look at the repo, the only thing changed in the "46 commits ahead of [official triton]" is the README. Somewhat sketchy.
- zorgmonkey 2y agoIt is mentioned at the top of the readme, but the actual code changes appear to be on the branch v3.1.x-windows (also a few other branches with -windows in the name). Also the triton seem willing to collaborate as long as the patches sent are reasonably minimal and high quality https://github.com/triton-lang/triton/pull/4045#issuecomment-2409922915 https://github.com/triton-lang/triton/pull/4045#issuecomment...
- lostmsu 2y agoThings might have changed since then, but I personally contributed a few Windows support-related changes back in 2021 as an independent contributor: https://github.com/triton-lang/triton/pulls?q=is%3Apr+is%3Aclosed+author%3Alostmsu https://github.com/triton-lang/triton/pulls?q=is%3Apr+is%3Ac...