3 ms·
(the following is speculation) text like hands belong to class of imagery satisfying two characteristics: 1) They are intricately structured, having many subco
by disconcision 3y ago
(the following is speculation)
text like hands belong to class of imagery satisfying two characteristics:
1) They are intricately structured, having many subcomponents which have precise spatial interrelationships over a range of scales; there are a lot of ways to make things that are like text/hands except wrong
2) The average person is intimately familiar with said structures, having spent thousands of hours engaging looking at them while performing complex tasks involving a visiospatial feedback loop.
image generation models tend to have trouble with (1), but people only tend to notice it when paired with (2).
(1) can be improved by scale and more balanced training data; consider that for a person, their own hands are very frequently in their own field of view, but the photos they take only rarely feature hands as the focus. this creates a differential bias.
as for (2), image models tend to generate all kinds of implausibilities that the average person doesn't notice. try generating a complex landscape and ask a geologist how it formed.
- deleted 3y ago[deleted]
- Wowfunhappy 3y ago> (1) [Hands] are intricately structured, having many subcomponents which have precise spatial interrelationships over a range of scales; there are a lot of ways to make things that are like text/hands except wrong 2) The average person is intimately familiar with said structures, having spent thousands of hours engaging looking at them while performing complex tasks involving a visiospatial feedback loop. Shouldn't this apply even more strongly to faces versus hands? AI seems to have a significantly easier time with those.
- wongarsu 3y agoFaces are probably vastly over-represented in the training data. Normal people and professional photographers alike love photographing faces, at zoom levels that are quite rare to experience in real life.
- dragonwriter 3y agoFaces don't have lots of repeating similar subcomponents beyond some things that are just two items in bilateral symmetry (teeth are a big exception, and teeth, when visible, can be a problem.) And, actually, faces still, especially outside of closeups of just the face, can be a problem, too, which is why a separate face restoration with a GAN or inpainting pass for faces with the same or different diffusion model is common.
- disconcision 3y agofaces are simpler since unlike hands most of their major constituents are at fixed relative positions to each other. but the flip side is that people are hyper-biased towards attending to facial details, hence why they were basically the first-handled special case