3 ms·
When running diffusion models locally you often generate a lot of garbage. People with eight fingers per hand, cats with four tails, etc. When that happens you
by braingenious 4y ago
When running diffusion models locally you often generate a lot of garbage. People with eight fingers per hand, cats with four tails, etc. When that happens you chalk that up to bad prompting or just the “magic” of image generation being too arcane to know.
When it comes to text, people don’t find “cats are an animal with four tails” an amusing statement in the same way that they do a drawing of one. The standard of acceptability is way higher.
- p-e-w 4y agoI'm not convinced by this explanation. Small language models (similar in size to the Stable Diffusion network) don't just produce incorrect statements like "cats are an animal with four tails", they produce incoherent sentences with no relation to the text they are supposed to extrapolate. It's not that they are wrong, they don't even make sense much of the time. That's not true for SD. Yes, many of its output images have flaws, but the overall image usually shows the desired subject, and roughly resembles something a human artist might paint.
- braingenious 4y ago> they produce incoherent sentences with no relation to the text they are supposed to extrapolate I think we’re agreeing here? Language models of the size of SD produce literally useless output. SD can produce kind of fun output that people can make use of sometimes?