3 ms·
Iirc we had good speedups on it with torch.compile, and I remember working on it. Let me see if I can find numbers…
by voz_ 3y ago
Iirc we had good speedups on it with torch.compile, and I remember working on it. Let me see if I can find numbers…
- brucethemoose2 3y agoIts about 20-40% depending on the GPU, from my tests. And only very recent builds of torch 2.1 (with dynamic input) work properly, and it still doesn't like certain input changes, or augmentations like controlnet. AIT is the most usable compiled implementation I have personally tested, but SHARK (running IREE/MLIR/Vulkan) and torch-mlir are said to be very good. Hidet is promising but doesn't really work yet. TVM doesn't have a complete implementation outside of the WebGPU demo.
- voz_ 3y agoTry head of master. If there’s any bugs or graph breaks you hit, lmk, I can take a look. My numbers say 71% with a few custom hacks. Glad the dynamic stuff is working out tho!
- brucethemoose2 3y agoI will, thanks! I have been away for a month, but I will start testing it again later and submit some issues I run into.
- voz_ 3y agoMy username without underscore, at meta. Email me any bugs, I can help file them on GH and lend a hand fixing.