4 ms·
Some truly impressive results. I'll pick my usual point here when a fancy new (generative) model comes out, and I'm sure some of the other commenters have allud
by desideratum 6y ago
Some truly impressive results. I'll pick my usual point here when a fancy new (generative) model comes out, and I'm sure some of the other commenters have alluded to this. The examples shown are likely from a set of well-defined (read: lots of data, high bias) input classes for the model. What would be really interesting is how the model generalizes to /object concepts/ that have yet to be seen, and which have abstract relationships to the examples it has seen. Another commenter here mentioned "red square on green square" working, but "large cube on small cube", not working. Humans are able to infer and understand such abstract concepts with very few examples, and this is something AI isn't as close to as it might seem.
- deleted 6y ago[deleted]
- sendtown_expwy 6y agoIt seems unlikely the model has seen "baby daikon radishes in tutus walking dogs," or cubes made out of porcupine textures, or any other number of examples the post gives.
- m3at 6y agoIt might not have seen that specific combination, but finding an anthropomorphized radish sure is easier than I thought: type "大根アニメ" in your search engine and you'll find plenty of results
- ronsor 6y agoAt least for certain types of art, sites such as pixiv and danbooru are useful for training ML models: all the images on them are tagged and classified already.
- numpad0 6y agoImage search “大根 擬人化” do return similar results to the AI-generated pictures, e.g. 3rd from top[0] in my environment, but sparse. “大根アニメ” in text search actually gives me results about an old hobbyist anime production group[1], some TV anime[2] with the word in title...hmm Then I found these[3][4] in Videos tab. Apparently there’s a 10-20 year old manga/merch/anime franchise of walking and talking daikon radish characters. So the daikon part is already figured in the dataset. The AI picked up the prior art and combined it with the dog part, which is still tremendous but maybe not “figuring out the daikon walking part on its own” tremendous. (btw anyone knows how best to refer to anime art style in Japanese? It’s a bit of mystery to me) 0: https://images.app.goo.gl/LPwveUJPWHr6oK8Y8 https://images.app.goo.gl/LPwveUJPWHr6oK8Y8 1: https://ja.wikipedia.org/wiki/DAICON_FILM https://ja.wikipedia.org/wiki/DAICON_FILM 2: https://ja.wikipedia.org/wiki/%E7%B7%B4%E9%A6%AC%E5%A4%A7%E6%A0%B9%E3%83%96%E3%83%A9%E3%82%B6%E3%83%BC%E3%82%BA https://ja.wikipedia.org/wiki/%E7%B7%B4%E9%A6%AC%E5%A4%A7%E6... 3: https://youtube.com/watch?v=J1vvut5DvSY https://youtube.com/watch?v=J1vvut5DvSY 4: https://youtu.be/1Gzu2lJuVDQ?t=42 https://youtu.be/1Gzu2lJuVDQ?t=42
- tkgally 6y ago> anyone knows how best to refer to anime art style in Japanese? The term mangachikku (漫画チック, マンガチック, "manga-tic") is sometimes used to refer to the art style typical of manga and anime; it can also refer to exaggerated, caricatured depictions in general. Perhaps anime fū irasuto (アニメ風イラスト, anime-style illustration), while a less colorful expression, would be closer to what you're looking for.
- Alex3917 6y agoIf you type in different plants and animals into GIS, you don’t even get the right species half the time. If GPT-3 has solved this problem, that would be substantially more impressive than drawing the images.
- adsche 6y agoYes, I don't really see impressive language (i.e. GPT3) results here? It seems to morph the images of the nouns in the prompt in an aesthetically-pleasing and almost artifact-free way (very cool!). But it does not seem 'understand' anything like some other commenters have said. Try '4 glasses on a table' and you will rarely see 4 glasses, even though that is a very well-defined input. I would be more impressed about the language model if it had a working prompt like: "A teapot that does not look like the image prompt." I think some of these examples trigger some kind of bias, where we think: "Oh wow, that armchair does look like an avocado!" - But morphing an armchair and an avocado will almost always look like both because they have similar shapes. And it does not 'understand' what you called 'object concepts', otherwise it should not produce armchairs where you clearly cannot sit in due to the avocado stone (or stem in the flower-related 'armchairs').
- ralfd 6y ago> I would be slightly more impressed about the language model if it had a working prompt like: "A teapot that does not look like the image prompt." Slightly? Jesus, you guys are hard to please.
- adsche 6y agoRight, that was unnecessary and I edited it out. What I meant is that 'not' is in principal an easy keyword to implement 'conservatively'. But yes, having this in a language model has proven to be very hard. Edit: Can I ask, what do you find impressive about the language model?
- dash2 6y agoPerhaps the rest of the world is less blasé - rightly or wrongly. I do get reminded of this: https://www.youtube.com/watch?v=oTcAWN5R5-I https://www.youtube.com/watch?v=oTcAWN5R5-I when I read some comments. I mean... we are telling the computer "draw me a picture of XXX" and it's actually doing it. To me that's utterly incredible.
- 6y ago
- spyder 6y agoYea, with these kind of generative examples, they should always include the closest matches from the training set to see how much it just "copied".
- londons_explore 6y agoIt's very hard to define closest...
- jonesn11 6y agoThis is a spot on point. My prediction is that it wouldn't be able to. Given its difficulty to generate correct counts of glasses, it seems as though it still struggles with systematic generalization and compositionality. As a point of reference, cherrypicking aside, it could model obscure but probably well-defined baby daikon radish in tutu walking dog, but couldn't model red on green on blue cubes. Maybe more sequential perception, action, video data or system-2 like paradigm, but it remains to be seen.
- hanniabu 6y agoSounds like the perfect case for a new captcha system. Generate a random phrase to search an image for, show the user those results, ask them to select all images matching that description.