7 ms·
DALL-E 2 is a true "holy shit" moment for me. It's actually hard to believe it's real.
by CarrieLab 4y ago
DALL-E 2 is a true "holy shit" moment for me. It's actually hard to believe it's real.
- PheonixPharts 4y agoIt's interesting because for me it's the opposite. I feel DALL-E 2 is the clear evidence that we're headed for the next AI-Winter soon. Don't get me wrong, DALL-E is remarkable. But it's remarkable in almost the exact same way ELIZA was remarkable in 1966, and Markov Chain generative models were in the 1990s. All of these, given the context computational powers at the time, were both miracles and parlor tricks at the same time. The trouble is that these demos are impressive, but not much has changed in how the world works. I work in AI/ML and every places I've seen the practical applications of these techniques has ranged the equivalent to adding sprinkles on a cake, to completely useless engineering nightmares that would ultimately create more value with their removal. Yea, a 3.5 billion parameter model is neat, but we know that in 1970 when we couldn't imagine training such a thing. The problem is that we've made essentially no progresses other than showing when you pour incredible human and physical resources into a pot the result looks cool. But when you do the accounting, and ask yourself "what has really changed with all these fantastic innovations in machine learning" the answer is surprisingly little. Dall-E 2 is the type of cool that will be figure 12-3 in a undergrad textbook in 20 years. Students will go "oh that's cool" and turn the page.
- nl 4y ago20 years ago we couldn't imagine doing visual search that worked in any way. Now we can do that in a model that runs locally on your phone. This is a pretty big deal.
- nahuel0x 4y agoI think you are totally, totally wrong. This is a turning point. Artificial Imagination is just going to revolutionize all content generation, arts and designing. It will expand to video/VR generation and the input prompts probably will came from realtime neurofeedback. And not to mention what fantastic tool will become to explore the humanity collective unconscious and our fundamental nature. In some way is reifying our hive mind.
- pavlov 4y agoThese models look backwards. They mine the past of what human imagination created and produce passable composites that lack individual expression. The more they’re used, the more obvious the technique becomes and soon we’ll be tired of it.
- ravi-delia 4y agoMy guess is that human creativity is mostly just technical skill + taste + random noise. DALL-E has the first one, and we could probably approximate the last one, so the middle is the only one that needs work. It feels like that's a similar issue to how GPT often ends up trailing off. Maybe some kind of improved attention would work? Or an improved version of the sampling trick?
- mlyle 4y agoHuman curation is where you get the taste from. As far as something to play with to get interesting ideas, and to take a first few cuts at implementing them, DALL-E is great. Then let's bring the human's technical skill and ability at curation into the mix.
- gfodor 4y agoThis was obviously the case for previous models. DALL-E 2 seems different - I’ve seen a few outputs that seem like genuinely novel creative artistic works.
- jazzyjackson 4y ago"Originality is the art of hiding your source" Harry Roolaart
- dgb23 4y agoThey are quite obviously similar to artistic work because they copy, mix and match stuff that artists have done.
- rndphs 4y agoI think the failures of people spouting hype and failing to deliver in ML has absolutely nothing to do with the real and immense progress which is happening in the field concurrently. I don't understand how one can look at GPT-3, DALL-E2, alpha go, alpha fold, etc and think hmmm... this is evidence of an AI winter. A balanced reading of the season imo suggests that we are in the brightest AI summer and there is no sign of even autumn coming. At least on the research side of things.
- 314 4y agoThe difference between the two views could be summarized in a textbook intro from twenty years ago: here is a list of problems that are not (now) AI. Back then it would have included chess, checkers and other games that were researched for their potential to lead to AI. In the end they all fell to specific methods that did not provide general progress. While the current progress on image related problems is great, if it does not lead to general advances then an AI winter will follow.
- spupe 4y agoI disagree. If we find a particular architecture is good for Chess, and another for image generation, then so be it. We would still have solved important problems. We are seeing both general and specific approaches improving rapidly. I don't think the AI winter was defined by a failure to reach AGI, but rather that they reached a Plateau and produced nothing of great commercial or even intellectual value for some years, while other computer science fields thrived. I would say the situation is the exact opposite right now.
- blinding-streak 4y agoCrawl, walk, run. You can't go directly from crawl to run. You need the intermediate steps (pun not intended)
- gwern 4y agoA very telling example, since we now have methods like Player of Games which apply a single general method to solve chess, checkers, ALE, DMLab-30, poker, Scotland Yard... And the diffusion models behind DALL-E apply to generative modeling of pretty much everything, whether audio or text or image or multimodal.
- gfodor 4y agoAlphaFold
- wyldfire 4y ago> I work in AI/ML and every places I've seen the practical applications of these techniques has ranged the equivalent to adding sprinkles on a cake, to completely useless engineering nightmares that would ultimately create more value with their removal. That "count the objects" app that can tell you how many items you have in a photo seems like a very practical application that wasn't possible with traditional CV before ML.
- kristianov 4y agoIt's only good for rough estimates. For accurate inventory you have to hand-count the items just to be sure. RFID tag is a much better technology.
- jazzyjackson 4y ago"count the objects" is an interesting example to choose, it brings to mind the 1962 analog computer called "numa-rete" , composed of a grid of photocell nuerons (physical circuits) that works in parallel to instantaneously count the objects sitting on its surface. Not far off from image recognition, save for the background segmentation ;) machine is briefly described here, I've seen a more thorough breakdown somewhere out there... https://distributedmuseum.illinois.edu/exhibit/biological_computer_laboratory/ https://distributedmuseum.illinois.edu/exhibit/biological_co...
- dx034 4y agoIf that worked well you could skip annual inventory, on which many companies spend huge amounts on. If you could just have a robot with video driving through the aisles of a supermarket to detect the number of items, misplaced items or expired items, that would be a huge value add and lead to less waste. But I doubt we're there yet.
- timClicks 4y agoAgree that we haven't really progressed past curve fitting. I'm hopeful that we'll see a resurgence in symbolic AI, rather than watching the whole domain freeze up for a few decades.
- IshKebab 4y agoHas symbolic AI ever produced anything that really demonstrates that it's the right approach? As far as I know DNNs are still way better at symbolic problems than symbolic AI. And calling it "curve fitting" is just disingenuous. There's quite a lot of evidence that "curve fitting" will scale for at least a few more orders of magnitude. Who knows, maybe the human brain is just "curve fitting".
- nl 4y ago> Has symbolic AI ever produced anything that really demonstrates that it's the right approach? No it hasn't. And the problems are clear - it's impossible to express anything in the rigid hierarchies that symbolic AI requires. Representing "Britain" in geographic, language, political and economic hierarchies does not allow a model to do any reasoning about what "British sense of humour" means. "Softer" structures that represent concepts as a "blob" in a multi-dimensional space is clearly a better approach (aka embeddings, and the even better representations that more complex models use are even better). Representing "Britian" as blob in a multidimensional space that is adjacent to concepts like "satire" and "surreal" as well as people like "John Cleese" lets a model reason about what "British sense of humour" means without being specifically trained. As GPT-J[1] says when prompted with "A good example of the British sense of humour": A good example of the British sense of humour is found in George Orwell’s novel, The Lion and the Unicorn. It’s a satire on socialism that is much more sophisticated than anything in contemporary Leftist intellectual thought. It was written during the war, in 1944. The story opens with a visit to a pub in a fictional village in England. The pub is named the Unicorn, but Orwell (or George, as he calls himself) has decided to call it the Lion. He explains: “The Lion is a pub, just like any other pub, where people drink in the evenings, and talk about their daily business and their hobbies and where they exchange ideas, views, points of view. But the difference is that nobody in the Lion ever argues about anything. They just sit there, saying nothing, drinking nothing, not even beer. People can come to the Lion and buy beer, and leave the Lion and not buy beer. Beer is freely on sale in the Lion, but nobody ever buys it. (One should note that George Orwell's "The Lion and the Unicorn" is nothing to do with a pub where no one buys beer. BUT the joke is kind of exactly like a British sense of humour). [1] https://textsynth.com/playground.html https://textsynth.com/playground.html
- stickfigure 4y agoELIZA and markov chain generators were curiosities without practical application. DALL-E 2 is useful now. It's easy to imagine it replacing 99designs, Getty Images, and most of the digital art services on fiverr. I can't imagine what is going to happen in the world of print-on-demand tshirts. And of course porn... what a market.
- kingcharles 4y agoMagazines and newspapers are regularly paying thousands of dollars for a single abstract illustration for an article. DALL-E 2 is going to replace a huge proportion of that. Not only the cost saving, but you can get an image minutes before you go to press.
- dx034 4y agoAnd what if the image is mostly comprised of an existing artwork with copyright? Not sure I'd take that legal risk as a newspaper.
- stickfigure 4y agoThen it will be rejected by the person using DALL-E, or someone else in the publishing chain? Maybe an automated reverse image search AI? I don't think anyone (yet) imagines all the humans at the NYT will be replaced by GPTX+DALL-E. We still have editors even though humans author the articles.
- dx034 4y agoDoes that exist? Can you determine how mcuch of a DALL-E picture is identical to copyrighted material? And what's the threshold anyway for copyright to apply? I'm not sure that's a process with set rules yet.
- whywhywhywhy 4y agoUntil it can deal with a round of feedback without completely changing the imagine then illustrators have nothing to worry about.
- visarga 4y agoThe first back-prop paper originates in 1970. Hindsight is 20/20, but back then it was not so clear what will follow, just like we couldn't have hoped to see such a model even a few years ago. About usefulness - the CLIP part of the model is a ready made zero shot image classifier. It reduces the amount of work needed for simple image classification tasks to just naming the classes. The generative part is good enough for illustrations. It will make an average web designer have the powers of a graphical artist. Unfortunately the models are restricted and expensive today. I hope to see a real open AI initiative to train such models and share the weights, but can't hope that from OpenAI.
- flycaliguy 4y agoIt’s been a few days and I still am stunned and still processing things. I’m a graphic artist for about 40% of my income.