4 ms·
I assume they released equivalents for Linux?
by 71a54xd 3y ago
I assume they released equivalents for Linux?
- OMGnotThatGuy 3y agoIt appears that it's already in there. The trick is that you have to install an extension in some apps to make it work. For instance, there's a TensorRT accelerator extension for Stable Diffusion WebUI that takes advantage of it. I'm installing it right now, though the docs leave a bit to be desired. https://github.com/NVIDIA/Stable-Diffusion-WebUI-TensorRT https://github.com/NVIDIA/Stable-Diffusion-WebUI-TensorRT
- OMGnotThatGuy 3y agoI got it installed and tested with an SD 1.5 model in Stable Diffusion Webui using a 4090. (SDXL models did not work.) I generated the same test prompt with and without their TensorRT acceleration. The generated images are nearly identical, with very minor, almost imperceptible differences, which is pretty standard for techniques that optimize the attention layers. Prompt: A cat wearing pajamas Model: deliberate_v2 Size: 768x512 Batch Size: 4 Clip Skip: 2 Samler: DPM++ 2S a Karras Steps: 70 Seed: 2024828515 Generation times avg over 5 runs ================================ without TensorRT: 19.2 seconds with TensorRT: 12.3 seconds Benchmarks Batch Size 1 / 2 / 4 ======================================= without TensorRT: 31.11 / 36.23 / 42.34 with TensorRT: 55.32 / 58.06 / 62.27 It's faster, but clearly it's not 4x faster. I suppose they cherrypicked benchmarks against generation techniques not using xformers or SDP Attention. Also, it appears to be limited to a max batch size of 4.
- gbertb 3y agowas wondering about this