5 ms·
Not sure if I'm missing a subtle nuance in your point but to me those "artifacts" are completely expected. Those artifacts like 3 arms are the patterns / output
by ffwd 4y ago
Not sure if I'm missing a subtle nuance in your point but to me those "artifacts" are completely expected. Those artifacts like 3 arms are the patterns / outputs in the model, but since it doesn't have a fundamental understanding of the patterns/objects like arms, it just blends many images of arms together and create things like 3 arms. Also why there are so many eyes, arms, legs and other things in other generative programs. It just spits out the training set in random configurations (ish).
I suspect also the reason the images look OK at a glance is because the images as a whole also represent patterns in the model so they actually come from "real life" / artist created images and thus have some sense of cohesion. But making the AI have all the right patterns so it never makes a mistake at all scales of the image while also being able to combine the pattern with real understanding of what they are conceptually is the real trick but until then it will be a "salad bowl collage" thing at random intervals.
The closest thing to the brain it looks like to me is simply the hierarchical nature of it which seems similar to v1/v2/the vision system in humans but I've only been told that, I'm no neuroscientist.
- l33tman 4y ago"It just spits out the training set in random configurations (ish)." is a pretty gross misrepresentation and oversimplification of how such a model works, akin to saying a human artist only spits out whatever they saw earlier in their life in random configurations, or saying that SD only spits out pixel values it has seen before, or combinations of pixel values that form edges, etc. FWIW I don't think there is anything particularly wrong in the model architectures or training data that in some fundamental way makes it impossible to always get 2 arms. After all, lots of other tricky things are almost always correct. I suspect it's a question of training time and model size mostly (not trivial of course as it's still expensive to re-train to check modified architectures etc). It's also a matter of diffusion sampling iterations and choice of sampler at inference time, for the case of SD.
- ffwd 4y agoI get your point, but I also think it depends on what you mean by oversimplification. Of course there is _a lot_ of stuff going on and things like SD capture all kinds of information, not just what I described, however, any way you want to describe it, capturing all the "constraints" and real life knowledge to perfectly create realistic images with all the details and all the higher abstractions correctly is not anywhere close I think. Also it's not only to always get 2 arms, it's to - at the same time - also get 2 ears, 2 eyes, perfect pupils, perfect fingers, perfect trees, perfect chairs, all simultaneously (if it is to be used at least in the mainstream) - etc you get my point. I also don't think there's anything wrong with the model architectures in themselves or the data, nor that it is impossible, only that it is hard and as you say I think it needs a lot of data and clever engineering to fix mistakes. It may even be possible to fix most mistakes, over time, which would be pretty impressive imo, but the absolute limits of what a model can produce/"contain" with our hardware is kind of an open question though interesting.
- cma 4y ago> but since it doesn't have a fundamental understanding of the patterns/objects like arms, it just blends many images of arms together and create things like 3 arms. But it rarely would put out say 8 arms. And the repeat artifacts are miles ahead of earlier stuff like clip draw or disco diffusion. So it does seem to have some idea of what's going on, just isn't perfect yet. It gets much worse without the 512x512 resolution, if you push both dimensions it loses scene coherence a lot more.
- ffwd 4y agoActually I should have mentioned this in the original post but I think the "3 arms" thing is kind of a bad example come to think of it. I think in general at least with SD, if's very unlikely to create 3 arms or or 8 arms if you for example ask for a person. Mostly it looks like a person because the text prompt maps to training data of persons, and so they will generally look like people with 2 arms. However, where it struggles I find is with finer details, and also _placement_ of things like arms, eyes, and relationships between them. This I think is because it only has a general idea of the shape of persons but no data for the exact specifics like where the arms, legs, eyes and so on should be placed in a very realistic anatomical way, and this is where I think the challenge is - the gap between a general pattern of a person and an extremely specific but also general one where it can modify it and transform it like a real human artist can. I'm not sure that's in the data exactly
- ouid 4y agothe hierarchical structure of the visual system is a completely emergent property of the fact that the visual brain is a three dimensional object encoding a 2 dimensional objecr efficiently by maintaining the spatial relation present in the data in the representation until you have finished using it. There's simply nowhere for the computation to go.
- nl 4y ago> it just blends many images of arms together and create things like 3 arms.. It just spits out the training set in random configurations (ish). This is a fundamental misunderstanding of what it is doing. You can see in work like https://twitter.com/lintool/status/1579830653126086656 https://twitter.com/lintool/status/1579830653126086656 that the model does have an understanding of what parts of the visual model represent as concepts.
- nullc 4y ago"it just blends many images of arms together and create things like 3 arms. Also why there are so many eyes, arms, legs and other things in other generative programs. It just spits out the training set in random configurations" That is thoroughly confused to the point of uselessness. The reason you get structural issues is because it's hard for the architecture to express large scale structure, but they get better and better at it simply by scaling up the network.
- ffwd 4y agoIt was poorly communicated in my post but what I was referring to there were earlier programs like in 2015 and not SD and newer ones. If you put in an image of a landscape it could fill out the landscape with eyes and elbows all over the generated image because it had no information or context for what an eye was or where it should go. But now you get SD, dalle and others which add more information not just by scaling, but also by mapping sentences/words to pre-existing images that already have cohesion. That way when you write in sentences to the text prompt, the model has more semantic information about what an eye is, but (IMO) only _indirectly_ because it will map a sentence to images that match that sentence. The question is always what information is actually contained in the training set and what is missing from it and when it creates an image where is the information from etc. In some ways, that means I think that meaning to us as humans, is different from scaling which is almost like pixel resolution except resolution of patterns and differentiation of patterns. Meaning in this sense is things like creating a doorway with no actual door, but still the doorway itself looks super realistic is rendered. You can fix it by scaling and increasing the differentiation of patterns I guess, but you can never fix all instances completely with scaling. That's why in some ways I think meaning is sort of orthogonal to scale, however on a philosophical level, they should converge but that's for another topic. I may have missed something in my thoughts here because this is sort of difficult to talk about without writing a book eventually.
- bawolff 4y agoFor me the distinction is that when an artist draws a 3 armed person it stands out immediately. This makes me feel like something is going on in the ai that is similar to our brains because the blindspots seem similar. And you're right that this is pretty unfounded intuition. Humans often seek meaning in things without meaning, so it might be unfounded. At some point all i can really do is shrug and say it feels "spooky" to me.
- wruza 4y agoIsn’t this because a training set usually consists 99% of implied things? Afaik, these never provide a full description like “…, also two hands, two legs, three fingers, arms not bent, adequately long limbs, leg asymmetry, cartoon physics, …”, and also never feed examples of a wrong geometry/biology/etc. I’m no NN guy, but to me all it seems as basically underconstrained and unrelated to “understanding”. It’s like these e.g. woodwork, magic trick, dancing, guitar, etc teachers who fail to message a way to do something and can only tell “look”, then just do it, ask you to repeat, and get annoyed when you fail again.