3 ms·
Have you tried torch.compile on the model with Torch 2.0? It should help a bunch!
by bufo 4y ago
Have you tried torch.compile on the model with Torch 2.0? It should help a bunch!
- tysam_and 4y agoVery great suggestion! Thanks for posting this! :D :)))) I'd like to, but colab seems to be super duper finnicky on the CUDA installations included. The version of Triton required for the default (and really--the fastest I think--) backend of Pytorch 2.0 for compilation requires a version of CUDA 1 step higher or so than what's on the Colab machines. I'm sure there's a fantastic way to sidestep that, I haven't just found it yet. Opening that up, should compile times not take as long, should _~hopefully~_ allow for a lot more cool things that are slow under an eager/interpreted mode -- like, perhaps, better parallel branches of execution, or some ops which technically are much faster due to being in-place, but were on-the-whole much slower due to requiring a lot of teeny tiny individual kernel launches. I'd estimate the network will improve maybe 1-2+ seconds with it (and I really hope the compile overhead isn't that long! D: D: D: D: That would/could be another nightmare in and of itself!)
- bufo 4y agoAh ok I might run some tests on my RTX A5000 and report back!