4 ms·
Their main claim of “faster” unfortunately is false. > Running on an Nvidia A100 GPU, Paella took 0.5 seconds to produce a 256x256-pixel image in eight steps,
by tehsauce 3y ago
Their main claim of “faster” unfortunately is false.
> Running on an Nvidia A100 GPU, Paella took 0.5 seconds to produce a 256x256-pixel image in eight steps, while Stable Diffusion took 3.2 seconds
Using the latest methods (torch 2.0 compile, improved schedulers) stable diffusion only takes about 1 second to generate a 512x512 image on an a100 gpu. A 256x256 image 1/4 the size presumably takes less than half that time.
So the corrected title is “Like diffusion but slightly slower and lower quality.”
- jeron 3y agoto add, there's a finetuned version of Stable Diffusion 1.5 that can output 5 fps for 256x256 (0.2 seconds per image)[0]. So over 2x faster than Paella at 256x256 [0]:https://www.reddit.com/r/StableDiffusion/comments/z3m97e/minisd_a_256x256_finetune_of_stable_diffusion_15/ https://www.reddit.com/r/StableDiffusion/comments/z3m97e/min...
- dome271 3y agoHey, (one of the authors here). First of all the blog post is talking about the v1 of the paper, which was extremely fast, but not comparable to SD in any way. The v2 in arxiv is slower and does not achieve 0.5 seconds, but performs much better and closer to SD. So no doubt on this. But I just want to mention that you should not compare apples with oranges. Torch.compile also makes Paella much faster and using an optimized sampling pipeline it would always be faster than SD at 256x256 if you keep the conditions the same. Of course you could talk about distilling SD and then you can achieve maybe 1 step predictions etc. But you could probably do the same to Paella. I think it's important to stick with the main improvement from the paper that naive sampling can be done with much less steps when sticking to the original method, while being simple in its theory and implementation. But hey, way to go and improve on Paella in the future maybe :D
- tehsauce 3y agoHi! I really appreciate folks like you conducting and publishing real research. There have been a ton of companies recently which have been very rosily promoting their new models. My criticism was only to push back on overly optimistic marketing, and regret that some of it was directed at you. If you have a link to the v2 paper, would love to take a look!