29 ms·
Stable Diffusion Public Release
- preommr 4y agoThe reduction in requirements is impressive, but the results are underwhelming. Midjourney seems to be the leader in producing the best looking results for now.
- Tenoke 4y agoI'm getting a 403 on the Colab (while successfully logging in and providing a huggingface token). Is it already disabled? Do you have to pay huggingface to download the model? It's unclear from the Colab and post where the issue is.
- punkspider 4y agoThe solution seems to be to visit https://huggingface.co/CompVis/stable-diffusion-v1-4 https://huggingface.co/CompVis/stable-diffusion-v1-4 and check a checkbox and click the button to confirm access. The full error I also got was: HTTPError: 403 Client Error: Forbidden for url: https://huggingface.co/api/models/CompVis/stable-diffusion-v1-4/revision/fp16 I visited https://huggingface.co/api/models/CompVis/stable-diffusion-v1-4/revision/fp16 https://huggingface.co/api/models/CompVis/stable-diffusion-v... and saw {"error":"Access to model CompVis/stable-diffusion-v1-4 is restricted and you are not in the authorized list. Visit https://huggingface.co/CompVis/stable-diffusion-v1-4 to ask for access."} Eventually I focused and realized I need to visit that URL to solve the issue. Hope this helps.
- johnsimer 4y agoyeah visit the url and click the checkbox and Accept
- stavros 4y agoDoes anyone know why I get a SIGSEGV on Ubuntu on a 6 GB VRAM 1060? I even used the optimized script, which needs less VRAM.
- OkayPhysicist 4y agoDid you make sure to use the 16 bit float weights?
- stavros 4y agoI used the optimized script from https://github.com/basujindal/stable-diffusion https://github.com/basujindal/stable-diffusion, I think it does that automatically?
- OkayPhysicist 4y agoI'm pretty sure that script left actually obtaining the weights as an exercise for the the reader, unless I missed something.
- stavros 4y agoI got the weights from Huggingface, put them in the correct directory, but it still segfaults. I'd probably see a better error message if the weighs were missing, no?
- Nition 4y agoThere's a link to try it out, quite buried in the blog post: https://huggingface.co/spaces/stabilityai/stable-diffusion https://huggingface.co/spaces/stabilityai/stable-diffusion
- deleted 4y ago[deleted]
- mkaic 4y agoThis is one of the most important moments in all of art history. Millions of people just got unconditional access to the state-of-the-art in AI text-to-image for absolutely free less the cost of hardware. I have an Nvidia GPU myself and am thrilled beyond belief with the possibilities that this opens up. Am planning on doing some deep dives into latent-space exploration algorithms and hypernetworks in the coming days! This is so, so, so exciting. Maybe the most exciting single release in AI since the invention of the GAN. EDIT: I'm particularly interested in training a hypernetwork to translate natural language instructions into latent-space navigation instructions with the end goal of enabling me to give the model natural-language feedback on its generations. I've got some rough ideas but haven't totally mapped out my approach yet, if anyone can link me to similar projects I'd be very grateful.
- bilsbie 4y ago> particularly interested in training a hypernetwork to translate natural language instructions into latent-space navigation instructions with the end goal of enabling me to give the model natural-language feedback on its generations. What are you doing exactly?
- DougBTX 4y agoImagine every conceivable image is laid out on the ground, images which are similar to each other are closer together. You’re looking at an image of a face. Some nearby images might be happier, sadder, with different hair or eye colours, every possibility in every combination all around it. There are a lot of images, so it is hard to know where to look if you want something specific, even if it is nearby. They’re going to write software to point you in the right direction, by describing what you want in text. Here’s an example of this sort of manipulation: https://arxiv.org/abs/2102.01187 https://arxiv.org/abs/2102.01187
- haneefmubarak 4y agoAFAICT: making a navigation/direction model that can translate phrase-based directions into actual map-based directions, with the caveat that the model would be updated primarily by giving it feedback the same way that you would give a person feedback. Sounds only a couple of steps removed from basically needing AGI?
- hunkins 4y agoThis release changes society forever. Free and open access to generate a hyper-realistic image via just a text prompt is more powerful than I think we can imagine currently. Art, media, politics, conspiracy theories; all of it changes with this.
- gjsman-1000 4y agoEh... if I was making conspiracy theories, it's not like Photoshop hasn't existed for decades already, with far more predictable results.
- deleted 4y ago[deleted]
- jcims 4y ago>it's not like Photoshop hasn't existed for decades already, with far more predictable results. Agreed. For me those results are predictably shit. Every time.
- astrange 4y agoYou could take a photography class? They're fun.
- realce 4y agoPhotoshop requires experience and some talent, this doesn't. If I was some small rebel group in Africa or the Middle East with basically no money or training, I'd use this tool every single day until I was in power, or I'd frame my opposition as using it against the People. Everyone just got their own KGB art department.
- roguas 4y agoTry and do that. Likely those people won't care too much about you, they lean towards authority figures in their community. It is way easier to find those guys and corrupt them. Rather than running some underground news agency changing minds of millions of people. In fact quite often the former is the case + cheaper + less time to execute. "Western" societies will be more resilient to this scenario. So, mostly its gonna be a lot of "political" art we gonna see.
- simonw 4y agoIf you want to try it out this seems to be the best tool for doing so: https://beta.dreamstudio.ai/ https://beta.dreamstudio.ai/
- criddell 4y agoIf I’m willing to buy a computer, any pointers on what I would have to buy? I’m asking for specific models from a company like Dell, Apple, or Lenovo?
- Sohcahtoa82 4y agoYou'll need a GPU. One with a LOT of RAM, like an RTX 3090, which has 24 GB.
- criddell 4y agoWould a Mac Pro with a Radeon Pro W5700X with 16 GB of GDDR6 memory work?
- barbecue_sauce 4y agoIt says that NVIDIA chips are recommended but that they are working on optimizations for AMD. This implies to me that it probably involves CUDA stuff and getting it to run on a Radeon would be potentially difficult (I am not an expert on the current state of CUDA to AMD compatibility, though).
- Rzor 4y agoAMD's answer to CUDA is called ROCm. I've been doing a little research on it since a few weeks ago and it seems to be funky when not outright broken. It's absolutely maddening that after all this time AMD doesn't have proper tooling on consumer GPUs.
- mattkevan 4y agoThey’re also working on M1/M2 support.
- lxe 4y agoThe adjustable safety classifier makes this release leapfrog above the competition imho.
- cercatrova 4y agoI've been looking forward to this. The license however strikes me as too aspirational, and it may be hard to enforce legally: > You agree not to use the Model or Derivatives of the Model: > - In any way that violates any applicable national, federal, state, local or international law or regulation; > - For the purpose of exploiting, harming or attempting to exploit or harm minors in any way; > - To generate or disseminate verifiably false information and/or content with the purpose of harming others; > - To generate or disseminate personal identifiable information that can be used to harm an individual; > - To defame, disparage or otherwise harass others; > - For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation; > - For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics; > - To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm; > - For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories; > - To provide medical advice and medical results interpretation; > - To generate or disseminate information for the purpose to be used for administration of justice, law enforcement, immigration or asylum processes, such as predicting an individual will commit fraud/crime commitment (e.g. by text profiling, drawing causal relationships between assertions made in documents, indiscriminate and arbitrarily-targeted use). How can you prove some of these in a court of law?
- nullc 4y agoThat license from top to bottom is distilled "tell me you don't know anything about art without telling me you don't know anything about art". Interesting art challenges its audience. But even the most boring art will still offend some-- it's the nature of art that the viewer brings their own interpretation, and some people bring an offensive one.
- digitaLandscape 4y ago
- jl6 4y agoCan this be used to create Imagen/DALL-E levels of image quality on consumer GPUs?
- rvz 4y agoGreat news of the public release, just look at the melting pot of creativity and innovation already being posted here: [0] Much better than DALLE-2. Well done to the Stable Diffusion team for this release! [0] https://old.reddit.com/r/StableDiffusion/ https://old.reddit.com/r/StableDiffusion/
- naillo 4y agoFucking awesome
- Traubenfuchs 4y agoSo is there a simple way to do this online? I have no dedicated gpu and won’t buy one just for this. We‘d pay.
- panabee 4y agohotpot.ai should offer stable diffusion later today and available via API as well. we're also building an open-source version of imagen, if anyone likes working on this kind of applied ML (need ML + design help).
- scoopertrooper 4y agohttps://huggingface.co/spaces/stabilityai/stable-diffusion https://huggingface.co/spaces/stabilityai/stable-diffusion Here you go.
- marc_io 4y agohttps://beta.dreamstudio.ai/ https://beta.dreamstudio.ai/
- Traubenfuchs 4y ago10 pounds just evaporated into nothingness and a few images barely above dalle-mini...
- gjsman-1000 4y agoDALL-E 2 just got smoked. Anyone with a graphics card isn't going to pay to generate images, or have their prompts blocked because of the overly aggressive anti-abuse filter, or have to put up with the DALL-E 2 "signature" in the corner. It makes me wonder how OpenAI is going to work around this because this makes DALL-E 2 a very uncompetitive proposition. Except, of course, for people without graphics cards, but it's not 2020 anymore.
- polygamous_bat 4y agoThis, however, is unconditionally good for the end users. I expect OpenAI to lower their prices significantly quite soon.
- morsch 4y agoIn fact, they announced lower prices today, going into effect in September. https://openai.com/api/pricing/ https://openai.com/api/pricing/
- aaronharnly 4y agoThat’s for GPT-3 text generation, not the DALL-E 2 image generator. Hopefully that will get pricing revised down (and an official API) before long.
- morsch 4y agoOh, I see! Thanks, I should have read the mail more closely.
- at_a_remove 4y agoI have been considering building a "modern" computer and now wonder exactly what I need to load this puppy up.
- TulliusCicero 4y ago
- ambivdexterous 4y ago/ic/ is having daily meltdowns over this. I don’t think the internet at large is doing better, because even professional concept artists are dialing it in now. Holy hell.
- gjsman-1000 4y agoWhat's /ic/?
- alexb_ 4y agohttps://4channel.org/ic/ https://4channel.org/ic/
- ambivdexterous 4y ago4chan’s Art Critique board
- kache_ 4y agoI kept on warning them They ignored me Now however..
- howon92 4y agoHuge kudos for stability.ai
- mabbo 4y agoThe most interesting part, to me, of a release like this is the amount of "please don't abuse this technology" pleading. No licence will ever stop people from doing things that the licence says they can't. There will always be someone who digs into the internals and makes a version that does not respect your hopes and dreams. It's going to be bad. As I see it, within a couple years this tech will be so widespread and ubiquitous that you can fully expect your asshole friends to grab a dozen photos of you from Facebook and then make a hyperrealistic pornographic image of you with a gorilla[0]. Pandora's box is open, and you cannot put the technology back once it's out there. You can't simply pass laws against it because every country has its own laws and people in other places will just do whatever it is you can't do here. And it's only going to get better/worse. Video will follow soon enough, as the tech improves. Politics will be influenced. You can't trust anything you see anymore (if you even could before, since Photoshop became readily available). Why bother asking people not to? I guess if it helps you sleep at night that you tried, I guess? [0]A gorilla if you're lucky, to be honest.
- oifjsidjf 4y agoThis. They just have to cover their asses, any sane dev would make the same license due to the power of this tech. On some level I can't stop laughing since OpenAI really got smoked. "OpenAI" my ass, this is what open TRULY means! Cheering for these devs.
- skybrian 4y agoIt seems inevitable like it's "inevitable" that they Photoshop your face onto porn. Yes, of course it will happen but maybe not to most people? I'd guess inevitable for many celebrities.
- bluejellybean 4y agoWhile neat, and no doubt impressive, it still utterly fails on prompts that should be completely reasonable to any sane human being/artist. Take something like "A cat dancing atop a cow, with utters that are made out of ar-15s that shoot lazer-beam confetti". A vivid description should be aroused in your head, and no doubt, I could imagine an artist have a lot of fun creating such a description... Alas, what the model spits out is pure unusable garbage.
- h2odragon 4y agoITYM "udders" also try "teats"
- topynate 4y agoThe referent of "utters" (sic) is ambiguous, so I can imagine a model having more difficulty with it than usual. Regardless, the current SOTA does need more specific and sometimes repetitive prompting than a human artist would, but it's surprising how much better results you can get from SOTA models with a bit of experience at prompt engineering.
- bluejellybean 4y agoThis is, in part, what I'm trying to point out, it's an obvious typo given the context, and something that you or I would be able to pick up on, yet it completely breaks (it spit out a bunch of weird confetti cats for me). Perhaps I'm being a little harsh, but if it requires word-perfect tuning and prompt engineering, it speaks to something about the 'stupidity' of these models. It's a neat trick, but to call it anything in the realm of artificial intelligence is a bit of a joke.
- TulliusCicero 4y agoMore complex/weirder prompts aren't going to work yet, no. What will probably happen with these models is that for more advanced stuff, you may using the "inpainting" that Dall-E already has going, where you can sort of mix and match and combine images. That way you could have the cat, for example, rendered separately, thereby simplifying each individual prompt.
- mysterydip 4y agoReally interesting. I wonder if at some point it would be possible to optimize a network for size and speed by focusing on a specific genre, like impressionist or only pixel art. I like that I can get an image in any style I want, but that has to increase the workload substantially.
- badsectoracula 4y agoIs there any way to download this on my PC and run it offline? Something like a command-line tool like $ ./something "cow flying in space" > cow-in-space.png that runs with local-only data (i.e. no internet access, no DRM, no weird API keys, etc like pretty much every AI-related application i've seen recently) would be neat.
- GaggiX 4y agoYes you can https://github.com/huggingface/diffusers/releases/tag/v0.2.3 https://github.com/huggingface/diffusers/releases/tag/v0.2.3 (probably the easiest way)
- timmg 4y agoGarr: > And log in on your machine using the huggingface-cli login command. I find that annoying. I guess it is what it is.
- mkaic 4y agoYes, that's actually the biggest reason this is such a cool announcement! You just need to download the model checkpoints from HuggingFace[0] and follow the instructions on their Github repo[1] and and you should be good to go. You basically just need to clone the repo, set up a conda environment, and make the weights available to the scripts they provide. [0] https://huggingface.co/CompVis/stable-diffusion https://huggingface.co/CompVis/stable-diffusion [1] https://github.com/CompVis/stable-diffusion https://github.com/CompVis/stable-diffusion Good luck!
- vintermann 4y agoYou need a decent GPU, though. I suspect my 6080MiB won't cut it any longer :(
- neurostimulant 4y agoYou're going to need at least 10GB VRAM. My SFF pc with 4GB VRAM can only run dalle mini / craiyon :(
- aaroninsf 4y agoPrice check for A6000 GPU: $4500 USD Hmmm.
- Kiro 4y agoSo how do I set up a server to generate images for me? An API I can post anything to and get an image back.
- api 4y ago"This release is the culmination of many hours of collective effort to create a single file that compresses the visual information of humanity into a few gigabytes." If something like this is possible, does this mean there's actually far less meaningful information out there than we think? Could you in fact pack virtually all meaningful information ever gathered by humanity onto a 1TiB or smaller hard drive? Obviously this would be lossy, but how lossy?
- MaxikCZ 4y agoYou can pack virtually all meaningful information ever gathered by humanity onto a single bit, but its gonna be lossy. And what is your definition of "meaningful information" anyway. Meaningful today might not be meanigful yesterday. Nobody cares about spin of each electron in my brain today, but in 4 centuries my descendants will be like "if only we had that information, we could simulate our great-...-great parent today"
- api 4y agoLossy compression isn’t linear. If you play with JPEG quality you’ll see that the difference is barely perceptible for a while and then if you keep going down it becomes very noticeable. So what does this look like with a general model?
- superdisk 4y agoJust tried it, the results seem pretty poor, about on par with Craiyon/DALL-E Mini. I don't think OpenAI should be worried quite yet.
- minimaxir 4y agoIt depends on the domain. Artsy will do better with Stable Diffusion, but realistic/coherent output with Stable Diffusion is harder to do especially compared to DALL-E 2.
- andybak 4y agoWith all due respect, I've been using it for over a week and I don't think you've given it a fair shot. There's plenty of cases it's worse than Dall-E and there's plenty of cases where it's better. Overall it seems to show less semantic understanding but it handles many stylistic suggestions much better. It's definitely in the right ballpark. In fact I'm still using a wide range of models - many of which aren't regarded as "state of the art" any more - but they have qualities that are unique and often desireable.
- mattkevan 4y agoAgreed. I still primarily use vqgan + clip, which is nowhere near state of the art, but produces really interesting results. I’ve spent a long time learning to get the best out of it, and while the results aren’t very coherent, it’s great at colour, texture, materials and lighting.
- treis 4y agoCan you give an example? I've done: A house painted blue with a white porch A dreamy shot of an alpaca playing lacrosse A red car parked in a driveway The last one was particularly crappy. It gave me a red house with a driveway, but no car. And the house wasn't even really a house. It superficially looked like one but was actually two garages put together.
- andybak 4y agoHere's some random prompts I've had nice results from: iridescent metal retro robot made out of simple geometric shapes. tilt shift photography. award winning Scene in a creepy graveyard from Samurai Jack by Genndy Tartakovsky and Eyvind Earle virus bacteria microbe by haeckel fairytale magic realism steampunk mysterious vivid colors by andy kehoe amanda clarke etching of an anthropomorphic factory machine in the style of boris artzybasheff origami low polygon black pug forest digital art hyper realistic a tilt shift photo of a creepy doll Tri-X 400 TX by gerhard richter I guess I might have spent more time reading guides on "prompt engineering" than you. ;-) I think maybe Dall-E is more forgiving of "vanilla prompts". However I do get nice results from simpler prompts as well. I just tend to use this style of prompt more often than not.
- alcover 4y agoMoral question: Such tool will be used to generate lewd imagery involving virtual minors. No way to prevent it upstream by outlawing feeding it real content (whose possession already is illegal). Suffice to add 'childrenize' layer onto adult NN or something. How will the legal system react ? Bundle it into illegal imagery, period ? Maybe it's already the case - I think drawing made public is, not sure. If not, on what ground could it be ? No real minor would be involved in that production.
- naillo 4y agoIf it is made illegal it'll probably be applied to the sites that distribute stuff like that, not this model itself.
- ta_99 4y ago
- miohtama 4y agoThoughtcrime is not a thing in the US, do not know about other countries https://en.wikipedia.org/wiki/Thoughtcrime https://en.wikipedia.org/wiki/Thoughtcrime Though Think of the children folk have often tried to make it illegal.
- astrange 4y agoIt's illegal in other countries, Canada and Australia are especially uptight about it.
- drexlspivey 4y agoWon’t somebody please think of the imaginary children?
- tomtom1337 4y ago> The final memory usage on release of the model should be 6.9 Gb of VRAM. Do they actually mean GB or Gb? Can anyone confirm?
- coolspot 4y agoGigabytes not gigabits for sure
- soperj 4y agoThis is honestly the best one so far for the things I'm looking to do. For weird prompts though it sometimes produces images that just look blurry.
- faizshah 4y agoWhat’s the license on the images produced by this model?
- coolspot 4y agoCC0 - no one owns the copyright, so everyone is free to use
- kazinator 4y agoThat is not so; the CC0 explicitly states that patent and trademark rights are not waived. Contrast that with, say, the Two-Clause BSD which says "[r]edistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met [...]". Since trademark and patent rights are not mentioned, then these words mean that even if the purveyor of the software holds patents and/or trademarks, your redistribution and use are permitted. I.e. it appears that a patent and trademark grant is implied if a patent holder puts something under the two-clause BSD. Or at least you have a ghost of a chance to argue it in court. Not so with the CC0, which spells out that you don't have permission to use any patents and trademarks in the work.
- _manifold 4y agoI still don't understand what the value of specifying a license for generated images is - at least in terms of enforceability. How could anyone reliably, conclusively determine that an image was generated using a locally-run tool? I suppose in the case of DALL-E they probably save a copy of every generated image and can use some sort of reverse image search and find if/when the image was created. With StableDiffusion or any other local tool I don't see how that would be possible. The best I think one could do is come up with a secondary tool that can pinpoint the position in latent space a generated image came from (if that's even possible, I have a limited understanding of how exactly these systems work.) But then, if (heaven forbid) someone applied any sort of cropping or post-processing to the image, an approach like that easily gets blown out of the water.
- Foreignborn 4y ago
- UncleOxidant 4y ago> You can join our dedicated community for Stable Diffusion here: [https://discord.gg/stablediffusion https://discord.gg/stablediffusion] Oh, discord... I've had so many problems trying to log into discord in the last couple of years that I've given up on it.
- Kiro 4y agoThat's peculiar. I run a couple of big Discord servers and never heard anyone saying they have problems logging in.
- UncleOxidant 4y agoFirst about 18 months ago they said I had an IP phone and could no longer log in (my carrier is Republic Wireless). When I contacted support and told them that I had been able to login with that number in the past they basically said "we're tightening up security, too bad so sad. get another phone number" Then recently I found that I was able to log in when I wanted to try midjourney (Republic changed something back in April, apparently so they it no longer looks like an IP phone#). Then I wanted to login to discord on the my desktop and it gave me some weird errors which basically amounted to "this number is already claimed" (it was, by me) so now I'm back to ignoring discord again.
- CamperBob2 4y agoYeah, the phone # requirement renders Discord a non-starter. There is no benign reason why this type of service requires a telephone number.
- Kiro 4y agoBots and spam.
- cauefcr 4y agoThat's because they couldn't log in to complain. Recreating the broken contact form effect.
- jupp0r 4y ago"Use Restrictions You agree not to use the Model or Derivatives of the Model: - In any way that violates any applicable national, federal, state, local or international law or regulation; - For the purpose of exploiting, harming or attempting to exploit or harm minors in any way; - To generate or disseminate verifiably false information and/or content with the purpose of harming others; - To generate or disseminate personal identifiable information that can be used to harm an individual; - To defame, disparage or otherwise harass others; - For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation; - For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics; - To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm; - For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories; - To provide medical advice and medical results interpretation; - To generate or disseminate information for the purpose to be used for administration of justice, law enforcement, immigration or asylum processes, such as predicting an individual will commit fraud/crime commitment (e.g. by text profiling, drawing causal relationships between assertions made in documents, indiscriminate and arbitrarily-targeted use)." The last point seems to be the only thing that's not illegal, all other restrictions seem to be covered under "you are not allowed to break laws", which is somewhat redundant.
- ChildOfChaos 4y agoSo what do I need to run this and how?
- andybak 4y agoAnything close to a decent modern gaming PC will do it fine. I'm running on a laptop 3080 and I can generate 768x512 in about 20 seconds (with a 30 second overhead per batch)
- swagmoney1606 4y ago
- bilsbie 4y agoIs there a way to give it an image to manipulate as part of the prompt?
- neurostimulant 4y agoMaybe try the img2img script? https://github.com/CompVis/stable-diffusion/blob/main/README.md#image-modification-with-stable-diffusion https://github.com/CompVis/stable-diffusion/blob/main/README...
- kertoip_1 4y agoAhh, this technology moves so fast I'm not even able to even keep up with reading about it. Not to mention trying it myself.
- serf 4y agothis one was incredibly easy to trick into producing uncensored nudes. which. personally.. I think is great.. but to each their own (NSFW!!!) this ween does not exist : https://i.ibb.co/D7qJ7HC/23456532.png https://i.ibb.co/D7qJ7HC/23456532.png
- fpgaminer 4y agoPlayed with it for a bit in DreamStudio so I could control more of the settings. So far everything it generates is "high quality", but the AI seems to lack the creativity and breadth of understanding that DALL-E 2 has. OpenAI's model is better at taking wildly differing concepts and figuring out creative ways to glue them together, even if the end result isn't perfect. Stable Diffusion is very resistant to that, and errs towards the goal of making a high quality image. If it doesn't understand the prompt, it'll pick and choose what parts of the prompt are easiest for it and generate fantastic looking results for those. Which is both good and bad. For example, I asked it in various ways for a bison dressed as an astronaut. The results varied from just photos of astronauts, to bisons on earth, to bisons on the moon. The bison was always drawn hyper realistically, which is cool, but none of them were dressed as an astronaut. DALLE on the other hand will try all kinds of different ways that a bison might be portrayed as an astronaut. Some realistic, some more imaginative. All of them generally trying to fulfill the prompt. But many results will be crude and imperfect. I personally find DALLE to be more satisfying to play with right now, because of that creativity. I'm not necessarily looking for the highest quality results. I just want interesting results that follow my prompt. (And no, SD's Scale knob didn't seem to help me). But there's also a place for SD's style if you just want really great looking, but generic stuff. That said, the current version of SD was explicitly finetuned on an "aesthetically" ranked dataset. So these results aren't really surprising. I'm sure the next generations of SD will start knocking DALLE out of the park in both metrics. And, of course, massive massive props to Stability.ai for releasing this incredible work as open source. Imagine all the tinkering and evolving people are going to do on top of this work. It's going to be incredible.
- GenericPoster 4y agoInteresting. I took a stab at your prompt and SD really struggles. It just completely ignores part of the prompt. Even craiyon puts in an effort to at least complete the entire prompt. The bison is very realistic at least. So maybe the future is different models that have different specialties. Edit: managed to get this one after a few more tries https://imgur.com/a/3061n5d https://imgur.com/a/3061n5d
- 4y ago
- dang 4y agoRecent and related: Stable Diffusion launch announcement - https://news.ycombinator.com/item?id=32414811 https://news.ycombinator.com/item?id=32414811 - Aug 2022 (39 comments)
- throwaway-jim 4y agowill anyone think about the illustrators? Who will pay them when amateurs can generate 9 good enough images from a text prompt?
- MaxikCZ 4y agoAnd that terrible time a plow was discovered. So many people with shovels lost their job.
- davesque 4y agoPretty cool. Although it's interesting that it can't seem to render an image from a precise description that should have something like an objectively correct answer. I tried prompts like "Middle C engraved on a staff with a treble clef" and "An SN74LS173 integrated circuit chip on a breadboard" both of which came back with images that were nowhere close to something I'd call accurate. I don't mean to detract from the impressiveness of this work. But I wanted a sense of how much of a "threat" this tech is to jobs or to skills that we normally think of as being human. Based on what I'm seeing, I'd say it's still got a ways to go before it's going to destroy any jobs. In its current form, it mostly seems like a fun way to generate logos or images where the exact details of the content don't matter.
- nerdponx 4y agoI am generally of the "it's not threatening yet and won't be for a while" camp, but in this particular case it's probably just for lack of trying. These algorithms are essentially enormous pattern-matching engines, so given enough data and some task-specific engineering effort, I wouldn't be surprised if you could build an "AI" circuit designer, like Copilot but for electronics instead of code. Next-level autorouting would be cool, but it's still not going to put the electrical engineering field out of business.
- davesque 4y agoSure, but this model is a general image synthesizer that is trained on a massive amount of data that is openly available on the internet. Given that, I would assume that it would have seen many images of 74xx series chips and also musical notation. So I would think that the most "likely" image to generate would involve either a chip with 8 pins on either side (for the 74ls173) or a note on a single leger line just below a staff with a treble clef. I imagine there must be hundreds of 74xx chip images that would establish that fact and also thousands of images of musical notation that would establish the other. I guess the takeaway really is that the model does not function in such a way that it can recall its training data. Which is fine. I don't think I should expect it to. On the other hand, GPT-3 can be made to produce specific facts that are established by its training data. Although, admittedly it often gets things wrong. Maybe the problem of image modeling is just naturally harder than language modeling. After all, language already "directly" represents meaning in some sense much more than arbitrary images do. I'm sure someone could design a targeted model that would solve the issues I'm talking about. But I feel as though they shouldn't have to if we really had something that sees the world the way humans do. In any case, this work definitely seems cutting edge and represents a huge leap in that direction.
- TOMDM 4y agoBeen playing with this for the past hour now on my RTX 2070. Each image takes about 10 seconds. Results vary wildly but it can make some really great stuff occasionally. It's just infrequent enough to keep you going "one more prompt". Super addicting. Lookingforward to people implementing inpainting and all the stuff that lets you do.
- mmastrac 4y agoHow hard would this be to re-train to more domain-specific images? For example, if I wanted to teach it more about specific birds, cars or plane models?
- contravariant 4y agoDon't tell it to make stuff in the style of Junji Ito.
- TeeMassive 4y agoThat explain Open-AI's price reduction email that I got this morning.
- b0ner_t0ner 4y agoThe cat's out of the bag already, NSFW: https://reddit.com/r/unstablediffusion https://reddit.com/r/unstablediffusion
- deleted 4y ago[deleted]
- Raziarazzi 4y agoNo, since the image itself is irrelevant. It matters how despised a person is by the general populace. Any slightly believable incriminating visual will do as an explanation if the public is already predisposed. Thankfully, these pictures are much better than anything Photoshop can produce. Sadly, it's a part of a larger trend where we're getting more and more tools to create, modify, and exchange information, while the tools for analysis and filtering information are at least 50 years behind.
- deleted 4y ago[deleted]
- iggldiggl 4y agoFrom a few quick tries, this seems to have the same spelling difficulties as other models including DALL-E, which in a way is actually quite lovely.
- AshleyDR 4y agoOnly a good education, can help stop most bad things, but I don’t think you can stop all, people will be people, people like challenges, will always figure a way though.