3 ms·
Hey, (one of the authors here). First of all the blog post is talking about the v1 of the paper, which was extremely fast, but not comparable to SD in any way.
by dome271 3y ago
Hey, (one of the authors here). First of all the blog post is talking about the v1 of the paper, which was extremely fast, but not comparable to SD in any way. The v2 in arxiv is slower and does not achieve 0.5 seconds, but performs much better and closer to SD. So no doubt on this. But I just want to mention that you should not compare apples with oranges. Torch.compile also makes Paella much faster and using an optimized sampling pipeline it would always be faster than SD at 256x256 if you keep the conditions the same. Of course you could talk about distilling SD and then you can achieve maybe 1 step predictions etc. But you could probably do the same to Paella. I think it's important to stick with the main improvement from the paper that naive sampling can be done with much less steps when sticking to the original method, while being simple in its theory and implementation. But hey, way to go and improve on Paella in the future maybe :D
- tehsauce 3y agoHi! I really appreciate folks like you conducting and publishing real research. There have been a ton of companies recently which have been very rosily promoting their new models. My criticism was only to push back on overly optimistic marketing, and regret that some of it was directed at you. If you have a link to the v2 paper, would love to take a look!