3 ms·
I can't figure out if the author is clickbaiting, or honestly does not know that you still need the base diffusion model to use this? >Nvidia's new Perfusion i
by Jackson__ 3y ago
I can't figure out if the author is clickbaiting, or honestly does not know that you still need the base diffusion model to use this?
>Nvidia's new Perfusion image generator
Calls it an image generator, when it isn't.
>In the rapidly evolving landscape of AI art creation tools, Nvidia researchers have introduced an innovative new text-to-image personalization method called Perfusion.
A spark of understanding, that it is in fact not an image model on its own.
>But it’s not a million-dollar super heavyweight model like its competitors.
And immediately fails to comprehend what "personalization method" meant. Neither LoRa, nor textual inversion, and even dream booth are million dollar heavyweight competition.
>Despite its small size, it’s able to outperform leading AI art generators like Stability AI's Stable Diffusion v1.5, the newly released Stable Diffusion XL (SDXL), and MidJourney in terms of efficiency of specific editions.
Going to quote the the paper[0] for this one:
>Our approach extends these pre-trained models to portray personalized concepts. It is applied with Stable-Diffusion [Rombach et al. 2021],
So it is in fact based on Stable Diffusion 1.4/5, meaning that "Despite it's small size" should be something like "Despite it's (bigger than original weights) size.", as you still need the base weights to do anything.
Overall I'm still excited to see for myself how the technique compares to the other methods, should it ever be released.
[0] https://arxiv.org/abs/2305.01644 https://arxiv.org/abs/2305.01644