5 ms·
Why are these models so bad at hands?
by Method-X 4y ago
Why are these models so bad at hands?
- shawabawa3 4y agofor what it's worth, humans are also in general terrible at drawing hands - i think it's just a difficult problem
- roselan 4y agoAnother reason I saw was that models were trained on 512x512 "portrait" images including very few hands. Added to the inherent complexity of hands, this throw off their generation.
- cma 4y agoHumans seem terrible at it in very different ways, and definitely don't get as good at other parts before getting good at hands.
- ye-olde-sysrq 4y agoI'm a layman but the gist afaict is: These models don't understand relationships between objects in a scene, especially between distant objects. So they can't do hands for the same reason they can't get legs on a table right. They know roughly what a table and a table leg look like, but they don't understand that there needs to be 3-4 of them at least, and they need to be spaced so that the table sits level, and the perspective they should have as a result. So, I've seen tables where it kind of gets it right that the legs are in the corners but then as the table legs go down, the front ones are mysteriously behind something that ought to be under the table. And sometimes it kind of loses track of a table leg or two - they melt into the background. Very similar problem with hands. They need a very specific orientation and shape and the fingers all need to consistently point in the right direction, and typically the same direction (except for when they don't like with a pointed finger, etc). Curious as to how these models handle it so much better than prior generations. Is it something novel, or a specific hand-based fix they put it, or is it just "we made the model bigger"?
- sho_hn 4y agoIt still feels unintuitive to me that models aren't able to infer these concepts from the training data given how consistently the training data follows them. It's not like there will be a lot of examples of bad hands in there.
- nashashmi 4y agoOr maybe the right models have not been built yet or plugged in? Another commenter told me about openpose information which is an AI that detects human poses. If that neuron is plugged in, it might lead to more accurate numbers. Stable diffusion is trying to do this.
- nuc1e0n 4y agoThere will be now though.
- stavros 4y agoMaybe the problem is that the model can't count, and just knows that each finger has a 75% chance to have another finger next to it.
- bluejay2387 4y agoA hand isn't so much a 'thing' as it is a complex asymmetric relationship of multiple elements that have to be within certain ratios of each other to fairly tight tolerances. Humans are very sensitive to those ratios. It's a hard problem.
- xdennis 4y agoBut can't you say the same about faces (except for symmetry) and AI seems to only produce gorgeous women?
- 323 4y agoThe number 4 (palm fingers) is very precise. You can't have 3 or 5. But you can have a variable number of stripes in tiger coat for example. It's difficult for AI to pickup that they need exactly 4. The fingers themselves are also almost identical, but not really. If you learn a "platonic finger" it's not good enough, you should learn each finger individually. There is only so much you can spend on them, you got a million other things to learn. And the raters of the model are much more likely to penalize a bad face than some off details in a hand.