2 ms·
Unfortunately it's a little bit tricky today. The main issue is that we rely heavily on torch.compile + Triton for performance in this repo, and there isn't an
by chillee 3y ago
Unfortunately it's a little bit tricky today. The main issue is that we rely heavily on torch.compile + Triton for performance in this repo, and there isn't an Apple Silicon backend either for torch.compile or Triton.
For example, there's an AMD backend for Triton (and it's also integrated into torch.compile), which is why we can mostly do the same optimizations on Nvidia and AMD GPUs.
Ideally, there'd be an Apple Silicon backend for Triton, and then this repo would mostly work out of the box :)