3 ms·
Yeah, for sure. I think for deployment purposes, many times these model conversions are necessary (such as if you don't want to use Python). However, I do thin
by chillee 3y ago
Yeah, for sure. I think for deployment purposes, many times these model conversions are necessary (such as if you don't want to use Python).
However, I do think these model conversions are often a significant pain for users.
So, in some sense, the goal here is to show that the performance component and the "convert your model for deployment" component can be disentangled.
We also have work on allowing you to "export" an AOT-compiled version of your model with torch.compile, and that should allow you to deploy your models to run in other settings.
- andy99 3y agoThanks for the reply. "show that the performance component and the "convert your model for deployment" component can be disentangled" makes sense. Also, I liked the part of the article about torch.compile producing faster matrix-vector multiplication than cublas. I've seen the same thing on CPU, that it's way faster to just write and manually optimize a loop over a bunch of dot products than it is to use BLAS routines because of how simple the "matmul" actually is. I don't know how widely known that is.