5 ms·
Illusion Diffusion: Optical Illusions Using Stable Diffusion
- alanbernstein 4y agoThis is great. Those examples are not the best quality, but they're impressive. That prompted me to generate ambigrams with stable diffusion. The results looked odd, as ambigrams tend to, but the "text" was largely illegible. I wonder when the state of the art will be able to handle that request.
- suyash 4y agolove ambigrams and illusions, any description on how one can create their own ? Thanks!
- GrantS 4y agoSince the technical nugget is hidden in the code, the fun trick here is to alternate on odd and even steps between moving toward a duck in the latent space image and moving toward a rabbit in the 90-degree rotated version of the latent space image. (Normally you would feed the output of step n right back in as input to step n+1. That’s what is not happening as usual here.)
- kantrarian 4y agoambigrams are cool you can rotate the term you want to ambigram-ize and write it underneath the word and gently merge them together, "column by column" it's nice you ask because I recently saw a "youtube short" about the name Klint that nicely depicts the idea (volume warning) https://www.youtube.com/shorts/3I6rkpAQXmI https://www.youtube.com/shorts/3I6rkpAQXmI
- goldfeld 4y agoIt's odd that image AIs are not ready to overlay text. If you ask Dall-E or Midjourney also to say a few letters they do lots of nearest random neighbors by not just scrambling the idea of the word but also scribbling anything on top that it thinks looks remotely like writing but is not in any language. Maybe it's still developing the ability to read and maybe secretly creating a completely new script and lang.
- ebalit 4y agoIt's a side effect of the way the text input is represented before being used by the model. It doesn't get the text as a sequence of chars but as a sequence of tokens. This paper [1] shows that giving character-level awareness to the model can improve the "visual spelling". 1: https://arxiv.org/abs/2212.10562 https://arxiv.org/abs/2212.10562
- prolyxis 4y agoThis mostly just left me with a greater appreciation for the seamlessness of the original rabbit-duck illusion.
- frob 4y agoThis mostly just left me appreciating the genius of Looney Tunes. https://youtu.be/TiVhptVBns4 https://youtu.be/TiVhptVBns4
- goldfeld 4y agoQuite cool! and creative use of AI. Do you consider the project is all fleshed out or are there improvements that could be done?
- IIAOPSW 4y agoI've generally been disappointed by my prompts for optical illusions. I thought it would be better at it. An optical illusion is basically what happens when you relax the constraints on a graphical depiction, allowing objects to be connected in ways that are inconsistent with 3d geometry. The trick is that the inconsistency has to be global not local. Anywhere you zoom in on still looks like normal 3d space. I expected SD to be good at this, as a priori it never had a conception of how 3d space must look to begin with. Here's where it sucked. It seems to have learned the superficial aesthetic of an optical illusion or of "Escher" without learning the relevant component. It spits out things that either aren't optical illusions, or are just random disconnected spattering of geometrical inconsistencies without any overarching theme. A person made optical illusion will generally have a single main loop of impossibly connected objects, or at least some simple overall topology. The illusion is expected to exist on the global scale of the image, not as a weird pocket of a mostly normal image.
- jimmySixDOF 4y agoAlso See: 151 Illusions & Visual Phenomena with explanations by Michael Bach [1] https://news.ycombinator.com/item?id=25045392 https://news.ycombinator.com/item?id=25045392