4 ms·
But it wasn't, and it does make a difference. Dall-E really wants to draw top hats on people and not cats because the prompt is ambiguous and top hats are norma
by origin_path 4y ago
But it wasn't, and it does make a difference. Dall-E really wants to draw top hats on people and not cats because the prompt is ambiguous and top hats are normally seen on humans so it struggles to overcome that bias. Neither robots not cats wear top hats so it's an easier problem to get right.
But the real problem here is the refusal to do basic and normal things, like depict people. That's not normal - it's deeply weird and tells us a lot about what must be going on inside Google's ai research effort.
- _dain_ 4y ago>But the real problem here is the refusal to do basic and normal things, like depict people. That's not normal - it's deeply weird and tells us a lot about what must be going on inside Google's ai research effort. Google is fighting a secret war against the Loab demon race that lives inside the high dimensional vector spaces. They've recently made incursions into our reality via Stable Diffusion.
- adamsmith143 4y agoThe inability to draw realistic humans is indeed strange but the question at hand is compositionality and so drawing a Robot with a top had is indeed more impressive precisely because it's not likely to be in the training data and shows a deeper understanding of the prompt. Presumably the model could randomly regurgitate a person with a top hat on that was seen in it's training data but that's not at all likely with a robot as you yourself said.
- origin_path 4y agoIt's not an inability, it's a policy choice, which is why it's weird. The question is why does Google think this rule is a good idea. Imagen could surely draw very good humans if allowed to. Robot looking at a cat wearing a top hat appears to be easier than with a human for DALL-E too, judging from the comments on Alexander's article, because both objects are neutral with respect to top hats. But really the whole set of prompts is poorly chosen. The original challenge of arbitrary shapes in relative positions seems the best way to test understanding of grammar and object relationships, exactly to avoid the "humans wear top hats and cats never do" problem. A better set of prompts is important - in this Gary Marcus is correct - exactly because there's no point defining a specific prompt if later you'll decide you accept a totally different prompt. That kind of invalidates the point of betting on well specified challenges to begin with.