3 ms·
The model behind Stable Difussion works similar to pareidolia[1], recognising "shapes" on random noise following the prompt theme, and refining that "mental ima
by TuringTest 4y ago
The model behind Stable Difussion works similar to pareidolia[1], recognising "shapes" on random noise following the prompt theme, and refining that "mental image" until it generates something that matches a recognizable image.
In these animations it's easy to see it, as the shape recognition is not very stable. For example, in the last video when the prompt changes from "people" to "bears" you can see a backpack turning into a bear head, which then turns into an open-mouth bear head. And then, the arm of the person carring the backpack is also turned into more bear heads.
The next steps in the evolution of this technique should be in exploring the relation between noise and subjects, so that you can create variations of the same image maintaining the recognizable parts stable.
[1] https://en.wikipedia.org/wiki/Pareidolia https://en.wikipedia.org/wiki/Pareidolia
- synu 4y agoIt's a pretty compelling, trippy effect even if we don't have good ways to work around it yet. This is some amazing work that takes advantage of it: https://twitter.com/xsteenbrugge/status/1558508866463219712?s=20&t=5mJvVcQQ89AznS4Bml2nbQ https://twitter.com/xsteenbrugge/status/1558508866463219712?...
- googlryas 4y agoThis reminds me of part of the opening scene of Adaptation: https://youtu.be/6Geq3wVvaNE?t=170 https://youtu.be/6Geq3wVvaNE?t=170 It cuts off before the full montage but you should get the idea.
- fragmede 4y agoDall-E introduced outpainting yesterday: https://openai.com/blog/dall-e-introducing-outpainting/ https://openai.com/blog/dall-e-introducing-outpainting/
- midlightdenight 4y agoIt looks like DALLE is just performing inpainting on pieces of the image outside the original frame. The zoom effect is scaling the original frame larger while keeping the frame, then performing image to image generation on it, or using the newly defined image in the frame and sending it back through diffusion (at least that’s my guess). It’s a different process than what DALLE has since inpainting does not overwrite already generated pieces. Stable diffusion can also do inpainting. You can make sliding images this way. Slowly translating the image out of the frame and then filling in the blank space with inpainting.
- operator-name 4y agoThat's actually a pretty good analogy for diffusion models. Another commenter linked https://www.reddit.com/r/dalle2/comments/vnw3z9/a_house_in_the_middle_of_a_beautiful_lush_field/ https://www.reddit.com/r/dalle2/comments/vnw3z9/a_house_in_t..., showing how "flat stability" (zoom out / scrolling) is possible via in/out painting with masks. The linked technique zooms and doesn't use masks, hence the instability.