3 ms·
There are some obvious mistakes in tools like this. Such as: Human faces are wrong, writing is usually scrambled, fingers look weird etc... Do you know if we ne
by curiousssnake 4y ago
There are some obvious mistakes in tools like this. Such as: Human faces are wrong, writing is usually scrambled, fingers look weird etc... Do you know if we need to have a major breakthrough similar to what happened 6 months ago to fix this or could these be fixed with incremental improvements in current techniques / datasets?
- FairlyInvolved 4y agoI wouldn't necessarily say a major breakthrough as such but I do think some architectural change is needed. There are concepts in images that we don't rely on purely visual understanding for - like words, we have a language model that we use to rely on when we see text in images. I think we need the same thing in our models to reach the next level of capabilities by combining models across different domains. I don't know if this manifests as pre-training with a language model and then expanding and updating the tensors as part of image training, or some more complicated merger of the models. To learn logical concepts just from images seems entirely impractical, like we can't rely on having enough images such that models can understand words coherently as language. You could draw a picture of a sign that says "children crossing" not because you can understand and remember exactly what an image of such a sign would look like, but because you have an understand of English and the character set that would let you reproduce it. If you tried to learn to create the same sign in Arabic you'd either need to see a huge number of signs to learn from or (more likely) build a language model for Arabic. The kind of abstract understandings that we know we can train in language models just aren't learned by image transformers at this scale (or likely any practical scale). A language model could easily understand: "A red cube is stacked on top of a blue plate, a green pyramid is balanced on the red cube" and infer things like the position of the pyramid relative to the blue plate, image models quickly fall over with such examples. An interesting nascent (and hacky) example of the benefits of combining models is people are using language models like GPT-3 to create better prompts for image models.
- andrewshadura 4y agoOn the other hand, those are all things humans have troubles with when drawing, unless they have an extraordinary talent or a lot of experience.
- Nursie 4y ago> Human faces are wrong, writing is usually scrambled, fingers look weird I think a lot of this has been solved in DALL-E already. It's pretty good at right-looking faces, and fingers. Text not so much... but it does appear that whatever OpenAI are doing behind the scenes, that's getting better too.
- PeterisP 4y agoThe writing issue demonstrably has been solved without any breakthroughs by simply making a larger model (dall-E 2 vs the publicly available dall-e), the same appears to be for other main issues as well - it's just that the publicly available versions are based on smaller/weaker models than the state of art because they're significantly cheaper to run. Also, I seem to recall that at least some models deliberately harmed generation of human faces (e.g. by selection of training data) to draw away attention from the deepfake/fakenews usecases and the related ethical,political and PR issues; I would assume that if any of them wanted to actually try and make specifically faces look good, that would be purely a matter of some engineering work without any breakthroughs needed - I mean, we have evidence from face-specific models that the same technical architecture can do decent faces.
- ShamelessC 4y agoNeither DALLE version 1 or 2 was completely released. Further, DALLE2 definitely still has issues with generation of text, although latent diffusion can do an okay job of it.
- astrange 4y agoIt’s Imagen that can competently generate text, and their ablation studies show it gains the ability at a size somewhat above DALLE2’s, IIRC.
- ShamelessC 4y agoApologies I was referencing personal experience w.r.t. latent diffusion. https://replicate.com/laion-ai/erlich https://replicate.com/laion-ai/erlich It still has issues of course, but a lot better at spelling than DALLE2.
- lm28469 4y ago> similar to what happened 6 months ago The major breakthrough that happened 6 months ago is that someone put their api behind a website for people to play with