3 ms·
I got it installed and tested with an SD 1.5 model in Stable Diffusion Webui using a 4090. (SDXL models did not work.) I generated the same test prompt with and
by OMGnotThatGuy 3y ago
I got it installed and tested with an SD 1.5 model in Stable Diffusion Webui using a 4090. (SDXL models did not work.) I generated the same test prompt with and without their TensorRT acceleration. The generated images are nearly identical, with very minor, almost imperceptible differences, which is pretty standard for techniques that optimize the attention layers.
Prompt: A cat wearing pajamas
Model: deliberate_v2
Size: 768x512
Batch Size: 4
Clip Skip: 2
Samler: DPM++ 2S a Karras
Steps: 70
Seed: 2024828515
Generation times avg over 5 runs
================================
without TensorRT: 19.2 seconds
with TensorRT: 12.3 seconds
Benchmarks
Batch Size 1 / 2 / 4
=======================================
without TensorRT: 31.11 / 36.23 / 42.34
with TensorRT: 55.32 / 58.06 / 62.27
It's faster, but clearly it's not 4x faster. I suppose they cherrypicked benchmarks against generation techniques not using xformers or SDP Attention. Also, it appears to be limited to a max batch size of 4.