5 ms·
I find SD to be amazing technology, but it still (mostly) sucks at producing "intelligent" images. It basically fancy math that turns noise into images (from th
by coldcode 4y ago
I find SD to be amazing technology, but it still (mostly) sucks at producing "intelligent" images. It basically fancy math that turns noise into images (from the opposite it trained on) but still has no idea what it is producing. If you run it long enough you eventually get lucky and find a gem. I like to try "George Washington riding a Unicorn in Times Square"; I've so far never gotten anything a first year art student can draw. I wonder how long it will take before something more "AI" than "ML" will have an understanding even close to what a simple human brain can process.
In the meantime it's fun to play with it, plus I'd like to better understand the noise training process.
- minimaxir 4y agoWith SD, you have to use modifier quality/positional/artist keywords, as vanilla inputs give the model too much freedom.
- jw1224 4y ago> "George Washington riding a Unicorn in Times Square" The “secret” to Stable Diffusion (and other CLIP-based models) is being as descriptive as possible. This prompt, whilst easy for humans to imagine, actually has a whole lot of ambiguity baked in. How high is the unicorn flying? Is the unicorn even flying, or on the ground? How old is George Washington? What visual style is the image in? Is the image from the perspective of a pedestrian at ground level, or from up at skyscraper level? The more ambiguous the prompt, the less cohesive the image. To demonstrate, here’s 4 renders from your original prompt: https://imgur.com/a/Jo4qfOp https://imgur.com/a/Jo4qfOp And here’s 4 using the prompt “George Washington riding a unicorn in Times Square, cinematic composition, concept art, digital illustration, detailed”: https://imgur.com/a/lB36JqC https://imgur.com/a/lB36JqC Certainly not perfect, but for an additional 15 seconds of effort, far better.
- l33tman 4y agoThe reason you can't get the images you want from it is not because of the noise diffusion process (after all, this is probably the closest similarity to how a human gets a flash of creativity) but the lack of a large language model in SD - it was deliberately scaled down so the result could fit in consumer GPUs. DALLE-2 uses a much larger language model and you can explain more complicated concepts to it. Googles Imagen likewise (not released though). It's mostly a matter of scaling to get this better.
- astrange 4y agoIt's not just size but also model architecture. DALLE mini (craiyon.com) has the opposite priority because of its different architecture; you can enter a complex prompt and it will follow it, but it's much slower and the image quality is a lot worse. SD prefers to make aesthetic pictures over listening to everything you tell it. You can improve this in SD by raising cfg_scale at the cost of some weird "oversharpening" artifacts. Or, you can make a crappy image in DallE mini and use that as the img2img prompt with SD to make it prettier. The real sign it's lacking intelligence is, if you ask it a question it won't draw the answer, it'll just draw the question. Of course, they could fix that too, it's got a GPT in it, they just don't let it recurse…
- l33tman 4y agoYeah true, I like dalle-mini :) It did seem to understand the prompts better. The training set also affects it, as the guidance signal competes with the diffusion-model's priors it learned from the training set (the cfg_scale) and I've found situations where it seems the priors are just encoded too strong it seems - for example with very well-known celebs or objects it's difficult to make variations. I guess it's interesting that these issues are kind of reflected in humans as well.
- dr_dshiv 4y ago> I like to try "George Washington riding a Unicorn in Times Square"; I've so far never gotten anything a first year art student can draw. Why the hell would a first year art student draw that? Flunk their ass. God damn dumb ass prompts I have to deal with. —Stable Diffusion
- Psychoshy_bc1q 4y agoyou might try "Pony Diffusion" for that :-) https://huggingface.co/AstraliteHeart/pony-diffusion https://huggingface.co/AstraliteHeart/pony-diffusion
- pmoriarty 4y agoThere are some prompts which yield bad results, but many other for which the results are nothing short of stunning. I would not write off AI generators based on some poor results you got, but take a look at the best results others have gotten, then learn to use the tools yourself to do the same. As an artist who's been making art for many decades now, AI art generation systems just blow me away. The best images I've seen from them are far better than a lot of what many real human artists can do, and this technology is just in its infancy. I can't even imagine how good it'll be in another 5 or 10 years.
- deleted 4y ago[deleted]