69 ms·
Imagen, a text-to-image diffusion model
- Mo3 4y agoIs the source in public domain already?
- londons_explore 4y ago>Figure 2: Non-cherry picked Imagen samples Hooray! Non-cherry-picked samples should be the norm.
- braingenious 4y agoThis is super cool and I want to play with it.
- neolander 4y agoIt really does look better than DALL-E, at least from the images on the site. Hard to believe how quickly progress is being made to lucid dreaming while awake.
- deleted 4y ago[deleted]
- Jyaif 4y agoJesus Christ. Unlike DALL-E 2, it gets the details right. It also can generate text. The quality is insanely good. This is absolutely mental.
- not2b 4y agoYes, the posted results are really good, but since we can't play with it we don't know how much cherry picking has been done.
- addajones 4y agoThis is absolutely amazingly insane. Wow.
- james-redwood 4y agoMetacalculus, a mass forecasting site, has steadily brought forward the prediction date for a weakly general AI. Jaw-dropping advances like this, only increase my confidence in this prediction. "The future is now, old man." https://www.metaculus.com/questions/3479/date-weakly-general-ai-system-is-devised/ https://www.metaculus.com/questions/3479/date-weakly-general...
- sydthrowaway 4y agoHow can we prepare for this? This will result in mass social unrest.
- refulgentis 4y agoYou think so? I'm very high on the Kool-Aid, image generation and text transformation models are core parts of my workflow. (Midjourney, GPT-3) It's still an unruly 7 year old at best. Results need to be verified. Prompt engineering and a sense of creativity are core competencies.
- visarga 4y ago> Prompt engineering and a sense of creativity are core competencies. It's funny that people are also prompting each other. Parents, friends, teachers, doctors, priests, politicians, managers and marketers are all prompting (advising) us to trigger desired behaviour. Powerful stuff - having a large model and knowing how to prompt it.
- deleted 4y ago[deleted]
- aaaaaaaaaaab 4y agoStock up on guns, ammo, cigarettes, water filters, canned food, and toilet paper.
- boppo1 4y agoNah, learn Spanish and first-aid. Being able to fix people is more useful than having commodities that will make you a target.
- daenz 4y ago>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones and a tendency for images portraying different professions to align with Western gender stereotypes. Finally, even when we focus generations away from people, our preliminary analysis indicates Imagen encodes a range of social and cultural biases when generating images of activities, events, and objects. We aim to make progress on several of these open challenges and limitations in future work. Really sad that breakthrough technologies are going to be withheld due to our inability to cope with the results.
- alphabetting 4y agoThere is a contingent of AI activists who spend a ton of time on Twitter that would beat Google like a drum with help from the media if they put out something they deemed racist or biased.
- ceeplusplus 4y agoThe ironic part is that these "social and cultural biases" are purely from a Western, American lens. The people writing that paragraph are completely oblivious to the idea that there could be other cultures other than the Western American one. In attempting to prevent "encoding of social and cultural biases" they have encoded such biases themselves into their own research.
- andybak 4y agoGreat. Now even if I do get a Dall-E 2 invite I'll still feel like I'm missing out!
- rvnx 4y agoIt's always the same with AI research: "we have something amazing but you can't use it because it's too powerful and we think you are an idiot who cannot use your own judgement."
- 2bitencryption 4y agoI can understand the reasoning behind this, though. Dall-E had an entire news cycle (on tech-minded publications, that is) that showcased just how amazing it was. Millions* of people became aware that technology like Dall-E exists, before anyone could get their hands on it and abuse it. (*a guestimate, but surely a close one) One day soon, inevitably, everyone will have access to something 10x better than Imagen and Dall-E. So at least the public is slowly getting acclimated to it before the inevitable "theater-goers running from a projected image of a train approaching the camera" moment
- andybak 4y agoAs someone that spent an evening trying to generate images of Hitler Lego I think they have a point.
- deleted 4y ago[deleted]
- sexy_panda 4y agoWould I have to implement this myself, or is there something ready to run?
- UncleOxidant 4y agoI think implementing this yourself is likely not doable unless you have the computing resources of a Google, Amazon or Facebook.
- sexy_panda 4y agoIt seems like lucidrains is currently working on an implementation [1] of it. I would love it. [1] https://github.com/lucidrains/imagen-pytorch https://github.com/lucidrains/imagen-pytorch
- manchmalscott 4y agoThe big thing I’m noticing over DALL-E is that it seems to be better at relative positioning. In a MKBHD video about DALLE it would get the elements but not always in the right order. I know google curated some specific images but it seems to be doing a better job there.
- benwikler 4y agoTotally—Imagen seems better at composition and relative positioning and text, while DALL-E seems better at lighting, backgrounds, and general artistry.
- kossTKR 4y agoYeah Dall-e looks amazing, to a mysterious degree even with hints of humour and irony, while imagen images look cheap, one dimensional and quite ugly to be honest. Still amazing that we're at a point where that's the case, they're both incredible developments.
- visarga 4y agoInteresting discovery they made > We show that scaling the pretrained text encoder size is more important than scaling the diffusion model size. There seems to be an unexpected level of synergy between text and vision models. Can't wait to see what video and audio modalities will add to the mix.
- gwern 4y agoI think that's unsurprising. With DALL-E 1, for example, scaling the VAE (the image model generating the actual pixels) hits very fast diminishing returns, and all your compute goes into the 'text encoder' generating the token sequence. Particularly as you approach the point where the image quality itself is superb and people increasingly turn to attacking the semantics & control of the prompt to degrade the quality ("...The donkey is holding a rope on one end, the octopus is holding onto the other. The donkey holds the rope in its mouth. A cat is jumping over the rope..."). For that sort of thing, it's hard to see how simply beefing up the raw pixel-generating part will help much: if the input seed is incorrect and doesn't correctly encode a thumbnail sketch of how all these animals ought to be engaging in outdoors sports, there's nothing some low-level pixel-munging neurons can do to help much.
- visarga 4y agoI was thinking more about our traditional ResNet50 trained on ImageNet vs CLIP. ResNet was limited to a thousand classes and brittle. CLIP can generalise to new concept combinations with ease. That changes the game, and the jump is based on NLP.
- ravi-delia 4y agoBasically makes sense, no? DALLE-2 suffered from misunderstanding propositional logic, treating prompts as less structured then it should have. That's a text model issue! Compared to that, scaling up the image isn't as important (especially with a few passes).
- espadrine 4y agoIs there a way to confirm that this extra processing relates to the language structure, and not the processing of concepts? I wouldn’t be surprised if the lack of video and 3D understanding in the image dataset training fails to understand things like the fear of heights, and the concept of gravity ends up being learned in the text processing weights.
- endisneigh 4y agoI give it a few years before Google makes stock images irrelevant.
- pphysch 4y agoThe entire "content" industry could get eaten by a few hundred people curating + touching-up output from these models.
- astrange 4y agoNo, competitive advantage means that it’s impossible to run out of jobs just because someone/something is better at it than you. (Consumer demand and boredom both being infinite is another thing working against it.)
- sydthrowaway 4y agoShort Getty images?
- tpmx 4y agoPrivately owned by the Getty family.
- tpmx 4y agoRolling this into Google Docs seems like a nobrainer.
- curiousgal 4y agoUntil they pull of the plug on it.
- semicolon_storm 4y agoOr rolling this into Google Image Search to create images that match users' search queries on the fly. Don't like any of the results from the real web? Well how about these we created just for you.
- armchairhacker 4y agoDoes it do partial image reconstruction like DALL-E2? Where you cut out part of an existing image and the neural network can fill it back in. I believe this type of content generation will be the next big thing or at least one of them. But people will want some customization to make their pictures “unique” and fix AI’s lack of creativity and other various shortcomings. Plus edit out the remaining lapses in logic/object separation (which there are some even in the given examples). Still, being able to create arbitrary stock photos is really useful and i bet these will flood small / low-budget projects
- xnx 4y agoOpenAi really thought they had done something with DALL-E, then Google's all "hold my beer".
- dntrkv 4y agoOpenAI*
- FargaColora 4y agoThis looks incredible but I do notice that all the images are of a similar theme. Specifically there are no human figures.
- influxmoment 4y agoI believe DALLE and likely this model excluded images of people so it could not be misused
- FargaColora 4y agoInteresting, I had not understood that!
- benwikler 4y agoWould be fascinated to see the DALL-E output for the same prompts as the ones used in this paper. If you've got DALL-E access and can try a few, please put links as replies!
- qclibre22 4y agoSee the paper here : https://gweb-research-imagen.appspot.com/paper.pdf https://gweb-research-imagen.appspot.com/paper.pdf Section E : "Comparison to GLIDE and DALL-E 2"
- thorum 4y agoImagen seems better at capturing details/nuance from the prompt, but subjectively the DALLE-2 images feel more “real” to me. Not sure why. Something about the lighting?
- ravi-delia 4y agoThat feels about right. Imagen has a better text processing model, so it can tease apart the prompt, but DALLE has a rocking image part.
- joeycodes 4y agoPosting a few comparisons here. https://twitter.com/joeyliaw/status/1528856081476116480?s=21&t=ItidPm4Hq78Yo5KqpVHsdA https://twitter.com/joeyliaw/status/1528856081476116480?s=21...
- rg111 4y agoImagen seems more realistic where Dall-E2 is more feel-good. That is what I feel personally.
- joeycodes 4y agoI agree with you, but for me, Dall·E 2 feels good because 90% of the time I can keep hitting the generate button and massage the prompt until I get something inspirational, surprisingly, or visually pleasing. Without access to Imagen, it's impossible for me to compare how much of the "realistic feels" of its images is constrained by the taste of the cherry-pickers.
- jandrese 4y agoIs there a way to try this out? DALL-E2 also had amazing demos but the limitations became apparent once real people had a chance to run their own queries.
- wmfrov 4y agoLooks like no, "The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access."
- nomel 4y ago> the risks of unrestricted open-access What exactly is the risk?
- tpmx 4y agoReally unpleasant content being produced, obviously.
- jimmygrapes 4y agoA variation on the axiom "you cannot idiot proof something because there's always a bigger idiot"
- varenc 4y agoSee section 6 titled “Conclusions, Limitations and Societal Impact” in the research paper: https://gweb-research-imagen.appspot.com/paper.pdf https://gweb-research-imagen.appspot.com/paper.pdf One quote: > “On the other hand, generative methods can be leveraged for malicious purposes, including harassment and misinformation spread [20], and raise many concerns regarding social and cultural exclusion and bias [67, 62, 68]”
- userbinator 4y ago
- shannifin 4y agoNice to see another company making progress in the area. I'd love to see more examples of different artistic styles though, my favorite DALL-E images are the ones that look like drawings.
- fortran77 4y ago> At this time we have decided not to release code or a public demo. Oh well.
- y04nn 4y agoReally impressive. If we are able to generate such detailed images, is there anything similar for text to music? I would I though that it would be simpler to achieve than text to image.
- nomel 4y agoCompare the size of a raw image file to a raw music file, to get an idea of the complexity difference.
- redox99 4y agoOur language is much more effective at describing images than music.
- tomatowurst 4y agowhy stop at audio? the pinnacle of this would be text-to-videos, equally indistinguishable from real thing.
- burlesona 4y agoThe way things look when still is much easier to fake than the way things move. I would expect AI development to follow a similar path to digital media generally, as its following the increasing difficulty and space requirements of digitally representing said media: text < basic sounds < images < advanced audio < video. What’s more impressive to me is how far ahead text-to-speech is, but I think the explanation is straightforward (the accessibility value has motivated us to work on that for a lot longer).
- ml_basics 4y agoWhy is this seemingly official Google blog post on this random non-Google domain?
- dekhn 4y agoI'n not certain but I think it's prelease. The paper says the site should be at https://imagen.research.google/ https://imagen.research.google/ but that host doesn't respond
- mmh0000 4y agoYou mean one of Google's domains? # whois appspot.com [Querying whois.verisign-grs.com] [Redirected to whois.markmonitor.com] [Querying whois.markmonitor.com] [whois.markmonitor.com] Domain Name: appspot.com Registry Domain ID: 145702338_DOMAIN_COM-VRSN Registrar WHOIS Server: whois.markmonitor.com Registrar URL: http://www.markmonitor.com Updated Date: 2022-02-06T09:29:56+0000 Creation Date: 2005-03-10T02:27:55+0000 Registrar Registration Expiration Date: 2023-03-10T00:00:00+0000 Registrar: MarkMonitor, Inc. Registrar IANA ID: 292 Registrar Abuse Contact Email: abusecomplaints@markmonitor.com Registrar Abuse Contact Phone: +1.2086851750 Domain Status: clientUpdateProhibited (https://www.icann.org/epp#clientUpdateProhibited) Domain Status: clientTransferProhibited (https://www.icann.org/epp#clientTransferProhibited) Domain Status: clientDeleteProhibited (https://www.icann.org/epp#clientDeleteProhibited) Domain Status: serverUpdateProhibited (https://www.icann.org/epp#serverUpdateProhibited) Domain Status: serverTransferProhibited (https://www.icann.org/epp#serverTransferProhibited) Domain Status: serverDeleteProhibited (https://www.icann.org/epp#serverDeleteProhibited) Registrant Organization: Google LLC Registrant State/Province: CA Registrant Country: US Registrant Email: Select Request Email Form at https://domains.markmonitor.com/whois/appspot.com Admin Organization: Google LLC Admin State/Province: CA Admin Country: US Admin Email: Select Request Email Form at https://domains.markmonitor.com/whois/appspot.com Tech Organization: Google LLC Tech State/Province: CA Tech Country: US Tech Email: Select Request Email Form at https://domains.markmonitor.com/whois/appspot.com Name Server: ns4.google.com Name Server: ns3.google.com Name Server: ns2.google.com Name Server: ns1.google.com
- jefftk 4y ago
- ShakataGaNai 4y agoAll of these AI findings are cool in theory. But until its accessible to some decent amount of people/customers - its basically useless fluff. You can tell me those pictures are generated by an AI and I might believe it, but until real people can actually test it... it's easy enough to fake. This page isn't even the remotest bit legit by the URL, It looks nicely put together and that's about it. Could have easily put together this with a graphic designer to fake it. Let be clear, I'm not actually saying it's fake. Just that all of these new "cool" things are more or less theoretical if nothing is getting released.
- cellis 4y agoInference times are key. If it can't be produced within reasonable latency, then there will be no real world use case for it because it's simply too expensive to run inference at scale.
- theptip 4y agoThere are plenty of usecases for generating art/images where a latency of days or weeks would be competitive with the current state of the art. For example, corporate graphics design, logos, brand photography, etc. I really do think inference time is a red herring for the first generation of these models. Sure, the more transformative use-cases like real-time content generation to replace movies/games, but there is a lot of value to be created prior to that point.
- dougmwne 4y agoThere's been much prior work done to take these models down from datacenter size to single GPU size. Given continued work in that area and improving GPU performance it seems like it's just a matter of years before inference can be cheap and local for even the most impressive of generation.
- unholiness 4y agoCertificate is expired, anyone have a mirror?
- minimaxir 4y agoGenerating at 64x64px then upscaling it probably gives the model a substantial performance boost (training speed/convergence) than working at 256x256 or 1024x1024 like DALL-E 2. Perhaps that approach to AI-generated art is the future.
- the__alchemist 4y agoI'll be skeptical until I see it in action, vice pre-selected results.
- tomatowurst 4y agowhen will there be a "DALL-E for porn" ? or is this domain also claimed by Puritans and morality gate keepers? The most in demand text-to-image is use case is for porn.
- astrange 4y agoTrain it yourself. Danbooru is a publicly available explicit dataset.
- tomatowurst 4y agoThis is not something you can train on a regular AWS gpu-instance without racking up millions of dollars in bills to my knowledge. Dataset isn't an issue its a capex issue.
- astrange 4y agoIt’s possible from scratch on not as much personally owned hardware as you’d think but will take a long time, months maybe. Luckily, training from scratch will hopefully be obsoleted by fine-tuning - if someone else releases a generally capable model then you can turn that into another one for lower cost.
- alimov 4y agoWould it be bad to release this with a big warning and flashing gifs letting people know of the issues it has and note that they are working to resolve them / ask for feedback / mention difficulties related to resolving the issues they identified?
- marcodiego 4y agoOk. Now, how about the legality of it generating socially unacceptable images like child porn?
- faizshah 4y agoWhat's the best open source or pre-trained text to image model?
- spyremeown 4y agoJesus, this is so awesome. I think it’s the first AI that really makes me have that “wow” sensation.
- dr_dshiv 4y agoHow the fck are things advancing so fast? Is it about to level off …or extend to new domains? What’s a comparable set of technical advances?
- dqpb 4y agoThis video by Juergen Schmidhuber discusses the acceleration of AI progress: https://youtu.be/pGftUCTqaGg https://youtu.be/pGftUCTqaGg
- astrange 4y agoBigger model = better because a lot of performance at this task is memorization or the “lottery ticket hypothesis”. An impressive advance would be a small model that’s capable of working from an external memory rather than memorizing it.
- colinmhayes 4y agoI wondered why all the pictures at the top had sunglasses on, then I saw a couple with eyes. Still some work to do on this one.
- mistrial9 4y agoReading a relatively-recent Machine Learning paper from some elite source, and after multiple repititions of bragging and puffery, in the middle of the paper, the charts show that they had beaten the score of a high-ranking algorithm in their specific domain, moving the best consistant result from 86% accuracy to 88% accuracy, somewhere around there. My response was: they got a lot of attention within their world by beating the previous score, no matter how small the improvement was.. it was a "winner take all" competition against other teams close to them; the accuracy of less than 90% is really of questionable value in a lot of real world problems; it was an enormous amount of math and effort for this team to make that small improvement. What I see is semi-poverty mindset among very smart people who appear to be treated in a way such that the winners get promotion, and everyone else is fired. That this sort of analysis with ML is useful for massive data sets at scale, where 90% is a lot of accuracy, not at all for the small sets of real world, human-scale problems where each result may matter a lot. The amount of years of training that these researchers had to go through, to participate in this apparently ruthless environment, are certainly like a lottery ticket, if you are in fact in a game where everyone but the winner has to find a new line of work. I think their masters live in Redmond, if I recall.. not looking it up at the moment.
- gwern 4y agoWhat you're missing is that the performance on a pretext task like ImageNet top-1 will transfer outside ImageNet, and as you go further into the high score regime, often a small % can yield qualitatively better results because the underlying NN has to solve harder and harder problems, eliciting true solutions rather than a patchwork of heuristics. Nothing in a Transformer's perplexity in predicting the next token tells you that at some point it suddenly starts being able to write flawless literary style parodies, and this is why the computer art people become virtuosos of CLIP variants and are excited by new ones, because each one attacks concepts in slightly different ways and a 'small' benchmark increase may unlock some awesome new visual flourish that the model didn't get before.
- londons_explore 4y agoIf you worked in a hospital and you managed to increase the survival rate from 86% to 88%, you too would be a hero. Sure, it's only 2%, but if it's on a problem where everyone else has been trying to make that improvement for a long time, and that improvement means big economic or social gains, then it's worth it.
- davikr 4y agoInteresting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.
- minimaxir 4y agoGranted that's a selection bias: you likely won't hear about the cases where legit obscene output occurs. (the only notable case I've heard is the AI Dungeon incident)
- deleted 4y ago[deleted]
- toxicFork 4y agoWhat is the AI dungeon incident?
- lelandfe 4y agohttps://www.vice.com/en/article/93ywpp/text-adventure-game-community-in-chaos-over-moderators-reading-their-erotica https://www.vice.com/en/article/93ywpp/text-adventure-game-c... TL;DR generative story site creators employ human moderation after horny people inevitably use site to make gross porn; horny people using site to make regular porn justifiably freaked out Bring your popcorn
- toxicFork 4y agoAI is for porn
- forgingahead 4y agoTime to update this song? https://www.youtube.com/watch?v=j6eFNRKEROw https://www.youtube.com/watch?v=j6eFNRKEROw
- bergenty 4y agoPrimarily Indian origin authors on both the DALL-E and this research paper. Just found that impressive considering they make up 1% of the population in the US.
- throwaway743 4y agohttps://github.com/lucidrains/imagen-pytorch https://github.com/lucidrains/imagen-pytorch
- CobrastanJorji 4y agoIs this a joke?
- throwaway743 4y agoNo
- deleted 4y ago[deleted]
- w1nk 4y agoTo expand a bit for the grandparent, if you check out this authors other repos you'll notice they have a thing for implementing these papers (multiple DALLE-2 implementations for instance). You should expect to see an implementation there pretty quickly I'd guess.
- xtreme 4y agoNot to diminish their contribution but implementing the model is only one third of the battle. The rest is building the training dataset and training the model on a big computer.
- w1nk 4y agoYou're not wrong that the dataset and compute are important, and if you browse the author's previous work, you'll see there are datasets available. The reproduction of DALL-E 2 required a dataset of similar size to the one imagen was trained on (see: https://arxiv.org/abs/2111.02114 https://arxiv.org/abs/2111.02114). The harder part here will be getting access to the compute required, but again, the folks involved in this project have access to lots of resources (they've already trained models of this size). We'll likely see some trained checkpoints as soon as they're done converging.
- jonahbenton 4y agoI know that some monstrous majority of cognitive processing is visual, hence the attention these visually creative models are rightfully getting, but personally I am much more interested in auditory information and would love to see a promptable model for music. Was just listening to "Land Down Under" from Men At Work. Would love to be able to prompt for another artist I have liked: "Tricky playing Land Down Under." I know of various generative music projects, going back decades, and would appreciate pointers, but as far as I am aware we are still some ways from Imagen/Dalle for music?
- addandsubtract 4y agoI agree. How cool would it be to get an 8 min version of your favorite song? Or an instant DnB remix? Or 10 more songs in the style of your favorite album?
- jonahbenton 4y agoYeah. I particularly love covers and often can hear in my head X playing Y's song. Would love tools to experiment with that for real. In practice, my guess is that even though Dall-e level performance in music generation would be stunning and incredible, it would also be tiresome and predictable to consume on any extended basis. I mean- that's my reaction to Dall-e- I find the images astonishing and magical but can only look at them for limited periods of time. At these early stages in this new world the outputs of real individual brains are still more interesting. But having tools like this to facilitate creation and inspiration by those brains- would be so so cool.
- exac 4y agoYou can sort of do that with https://fairuseify.ml https://fairuseify.ml
- aembleton 4y agoI tried that site and the music sounds the same. I wonder if you can use this to bypass YouTube content ID check.
- SemanticStrengh 4y agoDoes it outperform DALL-E V2?
- SemanticStrengh 4y agoNote that there was a close model in 2021 ignored by all https://paperswithcode.com/sota/text-to-image-generation-on-coco https://paperswithcode.com/sota/text-to-image-generation-on-... (on this benchmark) Also what is the score of dalle v2?
- deleted 4y ago[deleted]
- SemanticStrengh 4y agoThis competitor might be better for respecting spatial prepositions and photorealism but on a quick look i find the images more uncanny. DALL-E has IMHO better camera POV/distance and is able to make artistic/dreamy/beautiful images. I haven't yet seen this Google model be competitive for art and uncaniness. However progress is great and I might be wrong.
- jeffbee 4y agoIs there anything at all, besides the training images and labels, that would stop this from generating a convincing response to "A surveillance camera image of Jared Kushner, Vladimir Putin, and Alexandria Ocasio-Cortez naked on a sofa. Jeffrey Epstein is nearby, snorting coke off the back of Elvis"?
- astrange 4y ago- The current examples aren’t convincing pictures of “a shiba inu playing a guitar”. - If you made that picture with actors or in MS Paint, politics boomers on Facebook wouldn’t care either way. They’d just start claiming it’s real if they like the message.
- hn_throwaway_99 4y agoAs someone who has a layman's understanding of neural networks, and who did some neural network programming ~20 years ago before the real explosion of the field, can someone point to some resources where I can get a better understanding about how this magic works? I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. Just looking for more information about how the software actually works, even if there are big chunks of it that are "this is beyond your understanding without taking some in-depth courses".
- londons_explore 4y agoFigure A.4 in the linked paper is a good high level overview of this model. Shame it was hidden away on page 19 in the appendix! Each box you see there has a section in the paper explaining it in more detail.
- hn_throwaway_99 4y agoUhh, yeah, I'm going to need much more of an ELI5 than that! Looking at Figure A.4, I understand (again, at a very high-level) the first step of "Frozen Text Encoder", and I have a decent understanding of the upsampling techniques used in the last 2 diffusion model steps, but the middle "Text-to-Image Diffusion Model" step that magically outputs a 64x64 pixel image of an actual golden retriever wearing an actual blue checkered beret and red-dotted turtleneck is where I go "WTF??".
- f38zf5vdt 4y agoA good explanation is here. https://www.youtube.com/watch?v=344w5h24-h8 https://www.youtube.com/watch?v=344w5h24-h8
- sinenomine 4y ago> but the middle "Text-to-Image Diffusion Model" step that magically outputs a 64x64 pixel image of an actual golden retriever wearing an actual blue checkered beret and red-dotted turtleneck is where I go "WTF??". It doesn't output it outright, it basically forms it slowly, finding and strengthening more and more finer-grained features among the dwindling noise, combining the learned associations of memorized convolutional texture primitives vs encoded text embeddings. In the limit of enough data the associations and primitives turn out composable enough to suffice for out-of-distribution benchmark scenes. When you have a high-quality encoder of your modality into a compressed vector representation, the rest is optimization over a sufficiently high-dimensional, plastic computational substrate (model): https://moultano.wordpress.com/2020/10/18/why-deep-learning-works-even-though-it-shouldnt/ https://moultano.wordpress.com/2020/10/18/why-deep-learning-... It works because it should. The next question is: "What are the implications?". Can we meaningfully represent every available modality in a single latent space, and freely interconvert composable gestalts like this https://files.catbox.moe/rmy40q.jpg https://files.catbox.moe/rmy40q.jpg ?
- ma2rten 4y agoI get the impression that maybe DALL-E 2 produces slightly more diverse images? Compare Figure 2 in this paper with Figures 18-20 in the DALL-E 2 paper.
- benreesman 4y agoI apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication. And short of FAIR? D2 lacrosse. There are exceptions to such a brash generalization, NVIDIA’s group comes to mind, but it’s a very good rule of thumb. Or your whole face the next time you are tempted to doze behind the wheel of a Tesla. There are two big reasons for this: - the talent wants to work with the other talent, and through a combination of foresight and deep pockets Google got that exponent on their side right around the time NVIDIA cards started breaking ImageNet. Winning the Hinton bidding war clinched it. - the current approach of “how many Falcon Heavy launches worth of TPU can I throw at the same basic masked attention with residual feedback and a cute Fourier coloring” inherently favors deep pockets, and obviously MSFT, sorry OpenAI has that, but deep pockets also non-linearly scale outcomes when you’ve got in-house hardware for multiply-mixed precision. Now clearly we’re nowhere close to Maxwell’s Demon on this stuff, and sooner or later some bright spark is going to break the logjam of needing 10-100MM in compute to squeeze a few points out of a language benchmark. But the incentives are weird here: who, exactly, does it serve for us plebs to be able to train these things from scratch?
- meowface 4y agoNot elitist at all; I highly appreciate this post. I know the basics of ML but otherwise am clueless when it comes to the true depths of this field and it's interesting to hear this perspective.
- benreesman 4y agoI used a lot of jargon and lingo and inside baseball in that post, it was intended for people who have deep background. But if you’re interested I’m happy to (attempt) answers to anything that was jargon: by virtue of HN my answers will be peer-reviewed in real time, and with only modest luck, a true expert might chime in.
- 4y ago
- davelondon 4y agoI'M SQUEEZING MY PAPER!
- ALittleLight 4y agoInteresting to me that this one can draw legible text. DALLE models seem to generate weird glyphs that only look like text. The examples they show here have perfectly legible characters and correct spelling. The difference between this and DALLE makes me suspicious / curious. I wish I could play with this model.
- Tehdasi 4y agoStill has the issue with screwing up mechanical objects. In their demo checkout the wheels on the skateboards, all over the place.
- sdenton4 4y agoFor comparison, most humans can't draw a bicycle: https://www.wired.com/2016/04/can-draw-bikes-memory-definitely-cant/ https://www.wired.com/2016/04/can-draw-bikes-memory-definite...
- dclowd9901 4y agoI blame it on the surprisingly structural cleverness of a bicycle. Opposing triangles probably isn’t the first thing most people think of when they think of a bicycle (vs two wheels and some handlebars)
- gwern 4y agoThey also can't draw pennies, the letter 'g' with the loop, and so on (https://www.gwern.net/docs/psychology/illusion-of-depth/index https://www.gwern.net/docs/psychology/illusion-of-depth/inde...). Bicycles may be clever, but the shallowness of mental representation is real.
- gpt5 4y agoI only see the problem for the paintings. If you choose a photo it's good. Could be a problem in the source data (i.e. paintings of mechanical objects are imperfect).
- syspec 4y agoI'm curious why all of these tools seem to be almost tailored toward making meme images? The kind of early 2010's, over the top description of something that's ridiculous
- benreesman 4y agoThese things can make any image you can define in terms of a corpus of other images. That was true at lower resolution five years ago. To the extent that they get used for making bored ape images or whatever meme du juor, it says much more about the kind of pictures people want to see. I personally find the weird deep dreaming dogs with spikes coming out of their heads more mathematically interesting, but I can understand why that doesn’t sell as well.
- TaylorPhebillo 4y agoMy hunch is that they aren't tailored toward ridiculous images exactly, but if they demonstrated "a woman sitting in a chair reading", it would be really hard to tell if the result was a small modification of an image in the training data. If they demonstrate "A snake made out of corn", I have less concern about the model having a very close training example.
- B1FF_PSUVM 4y agoAlso, almost 40 years ago, the name of a laser printer capable of 200 dpi. Almost there, the Apple Laserwriter nailed it at 300 dpi. Sometimes sneaked an issue of the "SF-Lovers Digest" in between code printouts.
- lxe 4y agoHey I also wrote a neural net that generates perfect images. Here's a static site about it. With images it definitely generated! Can you use it? Is there a source? Hah, of course not, because ethics!
- Veedrac 4y agoI thought I was doing well after not being overly surprised by DALL-E 2 or Gato. How am I still not calibrated on this stuff? I know I am meant to be the one who constantly argues that language models already have sophisticated semantic understanding, and that you don't need visual senses to learn grounded world knowledge of this sort, but come on, you don't get to just throw T5 in a multimodal model as-is and have it work better than multimodal transformers! VLM[1] at least added fine-tuned internal components. Good lord we are screwed. And yet somehow I bet even this isn't going to kill off the they're just statistical interpolators meme. [1] https://www.deepmind.com/blog/tackling-multiple-tasks-with-a-single-visual-language-model https://www.deepmind.com/blog/tackling-multiple-tasks-with-a...
- benreesman 4y agoIt’s just my opinion but I think the meme you’re talking about is deeply related to other branches of science and philosophy: ranging from the trust old saw about AI being anything a computer hasn’t done yet to deep meditations on the nature of consciousness. They’re all fundamentally anthropocentric: people argue until they are blue in the face about what “intelligent” means but it’s always implicit that what they really mean is “how much like me is this other thing”. Language models, even more so than the vision models that got them funded have empirically demonstrated that knowing the probability of two things being adjacent in some latent space is at the boundary indistinguishable from creating and understanding language. I think the burden is on the bright hominids with both a reflexive language model and a sex drive to explain their pre-Copernican, unique place in the theory of computation rather than vice versa. A lot of these problems just aren’t problems anymore if performance on tasks supersedes “consciousness” as the thing we’re studying.
- ravi-delia 4y agoI'd argue that there is probably at least one leap in terms of human-level writing which isn't just pure prediction. Humans write with intent, which is how we can maintain long run structure. I definitely write like GPT while I'm not paying attention, but with the executive on the task I outperform it. For all we know this is solvable with some small tweak to architecture, and I rather doubt that a model which has solved this problem need be conscious (though our own solution seems correlated with consciousness), but it is one more step.
- ComputerGuru 4y agoIt seems to have the same "adjectives bleed into everything problem" that Dall-E does. Their slider with examples at the top showed a prompt along the lines of "a chrome plated duck with a golden beak confronting a turtle in a forest" and the resulting image was perfect - except the turtle had a golden shell.
- codemonkey-zeta 4y agoProbably just a frontend coding mistake, and not an error in the model, but in the interactive example if you select: "A photo of a Shiba Inu dog Wearing a (sic) sunglasses And black leather jacket Playing guitar In a garden" The Shiba Inu is not playing a guitar.
- didgeoridoo 4y agoFound the QA tester.
- spekcular 4y agoAlso, no sunglasses in "A photo of a raccoon wearing sunglasses and a red shirt riding a bike in a garden," and a few similar prompts (e.g. surfing).
- astrange 4y agoThere are visible “alignment” issues in some of their examples still. The marble koala DJ in the paper doesn’t use several of the keywords. They have an example “horse riding an astronaut” that no model produces a correct image for. It’d be interesting if models could explain themselves or print the caption they understand you as saying.
- discmonkey 4y agoFor people complaining that they can't play with the model... I work at Google and I also can't play with the model :'(
- make3 4y agoI mean I don't know how that makes it any better from a reproducibility stand point lol
- arthurcolle 4y agoHow does that make you feel?
- quickthrower2 4y agoProbably like an employee
- karmasimida 4y agoI mean inference on this cost not small money. I don't think they would host this for fun then.
- hathym 4y agooff-topic: as a google employee do you have unlimited gce credits?
- octocop 4y agois your team/division hiring?
- Firmwarrior 4y agoEvery tech megacorp is always hiring people who can jump through the flaming code hoops just right
- interblag 4y agoI think they address some of the reasoning behind this pretty clearly in the write-up as well? > The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access. I can see the argument here. It would be super fun to test this model's ability to generate arbitrary images, but "arbitrary" also contains space for a lot of distasteful stuff. Add in this point: > While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset which is known to contain a wide range of inappropriate content including pornographic imagery, racist slurs, and harmful social stereotypes. Imagen relies on text encoders trained on uncurated web-scale data, and thus inherits the social biases and limitations of large language models. As such, there is a risk that Imagen has encoded harmful stereotypes and representations, which guides our decision to not release Imagen for public use without further safeguards in place. That said, I hope they're serious about the "framework for responsible externalization" part, both because it would be really fun to play with this model and because it would be interesting to test it outside of their hand-picked examples.
- SnowHill9902 4y agoCould there exist a quine for this?
- hahajk 4y agoOff topic, but this caught my attention: “In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access.” I work for a big org myself, and I’ve wondered what it is exactly that makes people in big orgs so bad at saying things.
- dougmwne 4y agoI think they were being careful not to be too quotable there on CNN.
- iuppiter 4y agoso cool
- qz_kb 4y agoI have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely available the internet will become harder to scrape for good data and models will start training on their own output. The internet will also probably get worse for humans since search results will be completely polluted with these "sort of realistic" images which can ultimately be spit out at breakneck speed by smashing words from a dictionary together...
- deleted 4y ago[deleted]
- gwern 4y agoI don't think it will be a big deal, for multiple different reasons: https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-e-2-can-and-cannot-do?commentId=CWKFyJYfgoZfP9955 https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-...
- JayStavis 4y agoHuh, I had never thought of that. Makes it seem like there's a small window of authenticity closing. The irony is that if you had a great discriminator to separate the wheat from the chaff, that it would probably make its way into the next model and would no longer be useful. My only recommendation is that OpenAI et al should be tagging metadata for all generated images as synthetic. That would be a really interesting tag for media file formats (would be much better native than metadata though) and probably useful across a lot of domains.
- joshspankit 4y agoThe OpenAI access agreement actually says that you must add (or keep?) a watermark on any generated images, so you’re in good company with that line of thinking.
- 4y ago
- rhacker 4y agoNext phase of all this: Image to 3d printable template files compatible with various market available printers. Print me a racoon in a leather jacket riding a skateboard.
- d--b 4y agoOne thing that no one predicted in AI development was how good it would become at some completely unexpected tasks while being not so great at the ones we supposed/hoped it would be good. AI was expected to grow like a child. Somehow blurting out things that would show some increasing understanding on a deep level but poor syntax. In fact we get the exact opposite. AI is creating texts that are syntaxically correct and very decently articulated and pictures that are insanely good. And these texts and images are created from a text prompt?! There is no way to interface with the model other than by freeform text. That is so weird to me. Yet it doesn’t feel intelligent at all at first. You can’t ask it to draw “a chess game with a puzzle where white mates in 4 moves”. Yet sometimes GPT makes very surprising inferences. And it starts to feel like there is something going on a deeper level. DeepMind’s AlphaXxx models are more in line with how I expected things to go. Software that gets good at expert tasks that we as humans are too limited to handle. Where it’s headed, we don’t know. But I bet it’s going to be difficult to tell the “intelligence” from the “varnish”
- satokausi 4y agoI doubt 99% of humans can draw a ”chess game with a puzzle where white mates in 4 moves”
- bambax 4y agoMaybe not draw, but we can do an image search for "chess puzzle mate in 4" which gives plenty of results: https://www.google.com/search?q=chess+puzzle+mate+in+4&tbm=isch https://www.google.com/search?q=chess+puzzle+mate+in+4&tbm=i... It would be surprising if AI couldn't do the same search and produce a realistic drawing out of any one of the result puzzles.
- astrange 4y agoThey can with computer assistance, and this AI sort of has that in that it’s both some “intelligence” and a whole lot of memorized internet, with the issue that we don’t know how to separate those things.
- stevage 4y ago
- planb 4y agoSeeing the artificial restrictions to this model as well as to DALL-E 2, I can't help but ask myself why the porn industry isn't driving its own research. Given the size of that industry and the sheer abundance of training material, it seems just a matter of time until you can create photo realistic images of yourself with your favourite celebrity for a small fee. Is there anything I am missing? Can you only do this kind of research at google or openai scale?
- alexb_ 4y agoPorn is actually a really good litmus test to see if a money/media transfer technology has real promise. Pornography needs exactly 2 things to work well - a way to deliver media, and a way to collect money. If you truly have a system that can do one of those two things better than we currently can, and it's not just empty hype, it will be used for porn. "Empty hype" won't touch that stuff, but real-world usecases will. Unrelated to the main topic, but this is exactly why I think cryptocurrencies will only be used for illegal activities, or things you may want to hide, and nothing else. Because that's where it has found its usecase in porn.
- planb 4y agoWow. The last paragraph of my comment looked nearly identical to yours, but I deleted it before submitting because I didn't want to derail. Exactly my thoughts...
- rg111 4y agoTransfer learning is a thing. But I have not tried making generative models with out-of-distribution data before. Distributions other than main training data. There are several indie attempts that I am aware of. Mentioning them to the reply of this comment. (In case the comment gets deleted) The first layers should be general. But the later layers should not behave well to porn images. As they are more specialist layers learning distribution specific visual patters. Transfer learning is posssible.
- rg111 4y ago
- rishabhjain 4y agoUsed some of the same prompts and generated results with open source models, model I am using fails on long prompts but does well on short and descriptive prompts. Results: https://imgur.com/gallery/6qAK09o https://imgur.com/gallery/6qAK09o
- xyzal 4y agoTangentially related question: what is the best (~latest?) such a network uploaded to Colab one can toy with?
- awsrocks 4y ago
- wiz21c 4y agoI find it a bit disturbing that they talk about social impact of totally imaginary pictures of racoon. Of course, working in a golden lab at Google may twist your views on society.
- dougmwne 4y agoOh, I would say they are probably underestimating the impact. You only saw the images they thought couldn't raise alarm bells. Anyone will be able to create photorealistic images of anyone doing anything, Anything! This is certainly a dangerous and society altering tech. It won't all be teddy bears and racoons playing poker.
- octocop 4y agoWould be awesome to see a side by side comparison to DALL-E, generating from the same text
- mlfn 4y agoIt's in the PDF.
- deleted 4y ago[deleted]
- geonic 4y agoCan anybody give me short high-level explanation how the model achieves these results? I'm especially interested in the image synthesis, not the language parsing. For example, what kind of source images are used for the snake made of corn[0]? It's baffling to me how the corn is mapped to the snake body. [0] https://gweb-research-imagen.appspot.com/main_gallery_images/corn-snake-on-farm.jpg https://gweb-research-imagen.appspot.com/main_gallery_images...
- dave_sullivan 4y agoWell, first they parse the language into a high level vector representation. Then they take images and add noise and train a model to remove the noise so it can start with a noisy image and produce a clear image from it. Then they train a model to map from the word representation for text to the noisy image representation for the corresponding image. Then they upsample twice to get to good resolution. So text -> text representation -> most likely noised image space -> iteratively reduce noise N times -> upsample result Something like that, please correct anything I'm missing. Re: the snake corn question, it is mapping the "concept" of corn to the concept of a body as represented by intermediary learned vector representations.
- DougBTX 4y agoIn the paper they say about half the training data was an internal training set, and the other half came from: https://laion.ai/laion-400-open-dataset/ https://laion.ai/laion-400-open-dataset/
- kordlessagain 4y ago> Since guidance weights are used to control image quality and text alignment, we also report ablation results using curves that show the trade-off between CLIP and FID scores as a function of the guidance weights (see Fig. A.5a). We observe that larger variants of T5 encoder results in both better image-text alignment, and image fidelity. This emphasizes the effectiveness of large frozen text encoders for text-to-image models I usually consider myself fairly intelligent, but I know that when I read an AI research paper I'm going to feel dumb real quick. All I managed to extract from the paper was a) there isn't a clear explanation of how it's done that was written for lay people and b) they are concerned about the quality and biases in the training sets. Having thought about the problem of "building" an artificial means to visualize from thought, I have a very high level (dumb) view of this. Some human minds are capable of generating synthetic images from certain terms. If I say "visualize a GREEN apple sitting on a picnic table with a checkerboard table cloth", many people will create an image that approximately matches the query. They probably also see a red and white checkerboard cloth because that's what most people have trained their models on in the past. By leaving that part out of the query we can "see" biases "in the wild". Of course there are people that don't do generative in-mind imagery, but almost all of us do build some type of model in real time from our sensor inputs. That visual model is being continuously updated and is what is perceived by the mind "as being seen". Or, as the Gorillaz put it: … For me I say God, y'all can see me now 'Cos you don't see with your eye You perceive with your mind That's the end of it… To generatively produce strongly accurate imagery from text, a system needs enough reference material in the document collection. It needs to have sampled a lot of images of corn and snakes. It needs to be able to do image segmentation and probably perspective estimation. It needs a lot of semantic representations (optimized query of words) of what is being seen in a given image, across multiple "viewing models", even from humans (who also created/curated the collections). It needs to be able to "know" what corn looks like, even from the perspective of another model. It needs to know what "shape" a snake model takes and how combining the bitmask of the corn will affect perspective and framing of the final image. All of this information ends up inside the model's network. Miika Aittala at Nvidia Research has done several presentations on taking a model (imagined as a wireframe) and then mapping a bitmapped image onto it with a convolutional neural network. They have shown generative abilities for making brick walls that looks real, for example, from images of a bunch of brick walls and running those on various wireframes. Maybe Imagen is an example of the next step in this, by using diffusion models instead of the CNN for the generator and adding in semantic text mappings while varying the language models weights (i.e. allowing the language model to more broadly use related semantics when processing what is seen in a generated image). I'm probably wrong about half that. Here's my cut on how I saw this working from a few years ago: https://storage.googleapis.com/mitta-public/generate.PNG https://storage.googleapis.com/mitta-public/generate.PNG Regardless of how it works, it's AMAZING that we are here now. Very exciting!
- anoncow 4y agoAnother ad for DALL-E.
- causi 4y agoOne thing I find particularly fascinating is that all the elements of the resulting image have a relatively cohesive art style.
- beeskneecaps 4y agoIt’s terrifying that all of these models are one colab notebook away from unleashing unlimited, disastrous imagery on the internet. At least some companies are starting to realize this and are not releasing the source code. However they always manage to write a scientific paper and blog post detailing the exact process to create the model, so it will eventually be recreated by a third party. Meanwhile, Nvidia sees no problem with yeeting stylegan and and models that allow real humans to be realistically turned into animated puppets in 3d space. The inevitable end result of these scientific achievements will be orders of magnitude worse than deepfakes. Oh, or a panda wearing sunglasses, in the desert, digital art.
- isx726552 4y agoI am absolutely terrified of all this for a different reason: all human professions (not just art) will soon be replaced by “good enough” AI, creating a world flooded with auto-generated junk and billions of people trapped permanently in slums, because you can’t compete with free, and no one can earn a living any longer. It’s an old fear for sure but it seems to be getting closer and closer every day, and yet most of the discussion around these things seems to be variations of “isn’t this cool?”
- orblivion 4y agoAnd then once you take the (probably trivial) step where the computers come up with the ideas for the images, these images won't be interesting anymore because we know a human didn't even make it. It won't be funny in the same way. "Oh that was clever" doesn't make sense anymore. We could reach a new level of jaded. (Also, hello readers from the year 2032 when all of these predictions sound silly.)
- kurthr 4y agoDon't forget the training data for those computer "ideas" will be "attention" and targeted at the most vulnerable 80% of the market. I'd hope that it makes them less fearful and angry, but nope... that drives attention. I wonder what combination of UFOs, satanic cults, and immigrant hoards it will be.
- butz 4y agoLooking at example pictures it seems that this model has trouble with putting sunglasses on a racoon.
- Reiden 4y agoWhat's the limiting factor for model replication by others? Amount of compute? Model architecture? Quality / Quantity of training data? Would really appreciate insights on the subjects
- oakhaven 4y agoThis would generate great music videos for Bob Dylan songs. I'm thinking Gates of Eden, ".. upon four legged forest cloud, the cowboy angel rides" :D