17 ms·
Show HN: New AI edits images based on text instructions
This works suprisingly well. Just give it instructions like "make it winter" or "remove the cars" and the photo is altered.
Here are some examples of transformations it can make:
Golden gate bridge: https://raw.githubusercontent.com/brycedrennan/imaginAIry/master/assets/gg-bridge-suprise.gif https://raw.githubusercontent.com/brycedrennan/imaginAIry/ma...
Girl with a pearl earring: https://raw.githubusercontent.com/brycedrennan/imaginAIry/master/assets/girl_with_a_pearl_earring_suprise.gif https://raw.githubusercontent.com/brycedrennan/imaginAIry/ma...
I integrated this new InstructPix2Pix model into imaginAIry (python library) so it's easy to use for python developers.
- googie 4y agoHow to make it use my GPU (I have RTX 3070)? It complains about using sloooow CPU, but I don't see option to switch to GPU, which I think should be sufficient...? I'm running it on Windows 10.
- googie 4y agoNote, that I have the CUDA installed. Still the imaginAIry runs on CPU :(
- zepearl 4y agoThanks a lot!!! Works perfectly for me (Gentoo Linux + nVidia RTX3060 12GiB VRAM - I installed last week your package and it just worked, experimenting with it since then, telling about it parents & colleagues). The results (especially in relation to "people's faces") can vary a lot between ok/scary/great (I still have to understand how the options/parameters work), all in all it's a great package that's easy to handle & use. In general, if I don't specify a higher output resolution setting than the default (512x386 or something similar), with e.g. "-w 1024 -h 768", then faces get garbled/deformed like straight from a Stephen King novel => is this expected? Cheers :)
- deleted 4y ago[deleted]
- nullish_signal 4y ago>11GB VRAM Aaarrrgghh let me know when it's down to 4GB like Stable Diffusion The prompt-based masking sounds incredible, with either pixel +/- or Prompt Relevance +/- VERY impressive img2img capabilities!
- bryced 4y agoIt is stable diffusion but yes my fork does not have the memory optimizations needed to run it on only 4gb
- deleted 4y ago[deleted]
- do_anh_tu 4y agoOr, if there is a Colab version, I’d happy to pay Google for premium GPU.
- eega 4y agoWell, just open a new GPU Colab and create a cell mit „!pip install imaginairy“ and you should be good to go …
- bryced 4y agoIt does work in non-pro colab apparently. Here you go: https://colab.research.google.com/drive/1rOvQNs0Cmn_yU1bKWjCOHzGVDgZkaTtO?usp=sharing#scrollTo=P-h22zk43xiN https://colab.research.google.com/drive/1rOvQNs0Cmn_yU1bKWjC...
- NavinF 4y agoYou can get a used 2080Ti for under $300 on eBay
- smcleod 4y agoThat's a lot of money for most people. It also means they have to have a PC to put it in.
- deleted 4y ago[deleted]
- ilaksh 4y agoDoes anyone know if there is something like Google Cloud for GPUs but with an easy way to suspend the VM or container when not in use? Maybe I am just looking for container hosting with GPUs. I am just trying to avoid some of the basic VM admin stuff like creating, starting, stopping for SaaS if someone already has a way to do. Maybe this is something like what Elastic Beanstalk does.
- midlightdenight 4y agoMaybe not quite what you’re looking for, but I’ve seen some people mention banana.dev https://www.banana.dev/ https://www.banana.dev/ Never used it myself, but looks like AWS Lambda/GCP Cloud Functions tailored to ML models.
- TekMol 4y ago"Log in with Github". No thanks.
- punkspider 4y agoI think there's https://brev.dev https://brev.dev and https://banana.dev https://banana.dev
- naderkhalil 4y agoHey, founder of Brev.dev here. Brev lets you suspend the instances when not in use, and also auto-stops it after 3 hours of inactivity to avoid expensive surprises. Would love for you to give it a shot
- plufz 4y agoi’ve used https://www.genesiscloud.com/ https://www.genesiscloud.com/
- nicd 4y agovast.ai, paperspace.com
- PaulMest 4y agoI've played with several of these Stable Diffusion frameworks and followed many tutorials and imaginAIry fit my workflow the best. I actually wrote Bryce a thank you email in December after I made an advent calendar for my wife. Super excited to see continued development here to make this approachable to people who are familiar with Python, but don't want to deal with a lot of the overhead of building and configuring SD pipelines.
- bryced 4y agoThanks Paul!
- brycedriesenga 4y agoWhoa. Another Bryce D. So when do we fight?
- kennyadam 4y agoWhen the narwhal bacons of course. ugh.
- TekMol 4y agoHow can I try this? Can this be run on a Digitalocean VM? I looked around on DO's products, but none seems to advertise that it has a GPU. So maybe it is not possible?
- malux85 4y agoTry paperspace, they have GPUs and you can set billing limits to stop accidental overusage (no affiliation other than being a happy customer)
- bryced 4y agoHere is a google colab you can try it in: https://colab.research.google.com/drive/1rOvQNs0Cmn_yU1bKWjCOHzGVDgZkaTtO?usp=sharing https://colab.research.google.com/drive/1rOvQNs0Cmn_yU1bKWjC...
- patientplatypus 4y ago[dead]
- kumarm 4y agoDoesn't work if any people are in the photos: https://twitter.com/kumardexati/status/1616972740728356867/photo/1 https://twitter.com/kumardexati/status/1616972740728356867/p...
- sam1r 4y agoDid you have to tweet it wasn’t working versus just not making it a public “omg it’s not working it’s no good”
- bryced 4y agoWorks fine for me, you just need to adjust the strength of the edit.
- kumarm 4y agoYou mean steps?
- bryced 4y agono. in imaginairy it's called `--prompt-strength`. In other libraries it's called CFG or "classifier-free guidance". For the image edits I vary the strength of the effect from between 3-25
- bryced 4y agoFor the specific example you provide you could also use a prompt-based mask to prevent it from editing the person.
- lgas 4y agoIt does work on some things with people. I colorized a black and white photo of myself and then turned the colorized version into me as a Dwarven king.
- Gravyness 4y agoA similar tool: Instruct pix2pix to alter images by describing the changes required: https://huggingface.co/timbrooks/instruct-pix2pix#example https://huggingface.co/timbrooks/instruct-pix2pix#example Edit: Just noticed it is the same thing but wrapped, nevermind, pretty cool project!
- 0x4164 4y agoI hope there is a James Fridman version of this kind of AI.
- sam1r 4y agoThis is amazing! It’s only so long until video..
- pcrh 4y agoSee: https://news.ycombinator.com/item?id=34389041 https://news.ycombinator.com/item?id=34389041
- dr_kiszonka 4y agoVideo would be very useful too, but I expect running such models locally would be prohibitively expensive for most folks. (I am not talking about those $300k/year people here.)
- bobmaxup 4y agohttps://www.timothybrooks.com/instruct-pix2pix https://www.timothybrooks.com/instruct-pix2pix
- yieldcrv 4y ago“Add a dog in my arms” I’ll keep you posted how well this works for dating apps
- Uehreka 4y agoIt's a little premature, fine, but I want to start liquidating my rhetorical swaps here: I've been saying since last summer (sometimes on HN, sometimes elsewhere) that "prompt engineering" is BS and that in a world where AI gets better and better, expecting to develop lasting competency in an area of AI-adjacent performance (a.k.a. telling an AI what to do in exactly the right way to get the right result) is akin to expecting to develop a long-lasting business around hand-cranking people's cars for them when they fail to start. Like, come on. We're now seeing AIs take on tasks many people thought would never be doable by machine. And granted, many people (myself included to some extent) have adjusted their priors properly. And yet so many people act like AI is going to stall in its current lane and leave room for human work as opposed to developing orders of magnitudes better intelligence and obliterating all of its current flaws.
- chadcmulligan 4y agoHaven't we been here before? - see self driving cars.
- cududa 4y agoNo.
- c7b 4y agoLLMs and Image AIs are the opposite of self-driving cars. "Everybody" had concrete expectations for at least half a decade now that the moment where self-driving cars would surpass human ability was imminent, yet the tech hasn't lived up to it (yet). While practically nobody was expecting AI to be able to do the jobs of artists, programmers or poets anywhere near human level anytime soon, yet here we are.
- Der_Einzige 4y agoStill bad at poetry due to the tokenizer though. I wrote a whole paper on how to fix it: https://paperswithcode.com/paper/most-language-models-can-be-poets-too-an-ai https://paperswithcode.com/paper/most-language-models-can-be...
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- Der_Einzige 4y agoHoping that this is quickly implemented into the automatic1111 webUI.
- kewp 4y agoanyone know how to use this? kind of confusing install instructions in the readme
- bryced 4y agoIf you're used to installing python packages it should be relatively easy. There are other projects with nice UIs but that's not what this library is for.
- kadoban 4y agoIf you don't care what exact tool in particular, https://github.com/AUTOMATIC1111/stable-diffusion-webui https://github.com/AUTOMATIC1111/stable-diffusion-webui is the easiest to install I think and gives lots of toys to play with (txt2img, img2img, teach it your likeness, etc.)
- nicbou 4y agoCan it make it pop? Because that was the #1 request I remember dealing with.
- TekMol 4y ago#1 request of what, for what, requested by whom?
- stefanvdw1 4y agoThat is a common request when working with clients. They have a hard time describing what they want so end up asking to “make it pop”
- tamrix 4y agoMaybe but it could put their business logo anywhere!
- prox 4y agoI don’t know why people use this “AI” thing, I have been using make my logo bigger cream (tm) for ages with success. https://www.youtube.com/watch?v=GOwi3x92teo https://www.youtube.com/watch?v=GOwi3x92teo ;)
- marcosdumay 4y agoHum... Is that whitespace eliminator still on sale?
- c7b 4y agoMany thanks to the OP, can't wait to try this out! I have a question I'm hoping to slide in here: I remember there were also solutions for doing things like "take this character and now make it do various things". Does anyone remember what the general term for that was, and some solutions (pretty sure I've seen this on here, apparently forgot to bookmark). PS: I'm not trying to make a comic book, I'm trying to help a friend solve a far more basic business problem (trying to get clients to pay their bills on time).
- bryced 4y agodreambooth perhaps?
- c7b 4y agoDreambooth is what I'm using now, but I think I remember the concept had a specific name, something like 'context transfer' or so (pretty sure that was not the term) and tools that were pretty good at it that came out before Dreambooth. If I could at least remember the term it might be easier to search for them. Dreambooth is ok at it, but it requires multiple images (you often read 30, but I've actually had decent results with as few as five) to recognize what it's supposed to replicate. I remember there were tools that were more adapted to the workflow "create a humanoid cartoon character with a bunny face", pick one image that you like, and then "now show that same character in scene X, eg teaching in a classroom" or "wearing a cowboy outfit".
- airbreather 4y agoI'm getting mixed results, and for a given topic it seems to invariably give a better result first time you ask, then not so good if you ask again. It could be random and my imagination, but seems that way.
- natch 4y agoThe headline and the heavy promotional verbiage on the site seems to be claiming this is some new functionality we didn’t have before. Image2image with text instructions isn’t new as the headline implies. InvokeAI (and a few other projects as well) already does all this stuff much better unless I’m missing something. There are plenty of stable diffusion wrappers. Why not help improve them instead of copying them? I’m not against having enthusiasm for one’s project, but tell us why this is different and please don’t pretend the other projects don’t have this stuff.
- bryced 4y agoI'm not aware of any pre-existing open-source model that selectively edits images (leaving some parts untouched) based on instructions. This new method is much better than the image2image that shipped with the original stable diffusion. I'm looking at the InvokeAI docs right now and don't see anything like this feature. We previously had smart-masks, but InstructPix2Pix mostly does away with the need for those as well. If I am mistaken please provide links to these prior features.
- WesolyKubeczek 4y agoSo, Deckard can ask it to enhance, finally :)
- sebastiennight 4y agoIt's very interesting, thanks! I've noticed (on the Spock example) that "make him smile" didn't produce a very... "comely" result (he basically becomes a vampire). I was thinking of deploying something like that in one of our app features, but I'm scared of making our Users look like vampires :-) Is it your experience that the model struggles more with faces than with other changes?
- bryced 4y agoYes if you're not careful it can ruin the face. You can play with the strength factor to see if something can be worked out. Bigger faces are safer.
- sideshowb 4y agoIs there a link to how this works - in terms of nn architecture to combine the embedding of the existing image with the edit instruction?
- bryced 4y agohttps://arxiv.org/abs/2211.09800 https://arxiv.org/abs/2211.09800
- odedbend 4y agoWhere can I find more data about the work you did to create this?
- bryced 4y agoI did the work to wrap it up and make it "easy" to install. The researchers who did the real work can be found here: https://www.timothybrooks.com/instruct-pix2pix https://www.timothybrooks.com/instruct-pix2pix
- perfopt 4y agoHow does this work? When I run it on a machine with a GPU (pytorch, CUDA etc installed) I still see it downloading files for each prompt. Is the image being generated on the cloud somewhere or on my local machine? Why the downloads?
- bryced 4y agoShouldn't be downloads per prompt. Processing happens on your machine. It does download models as needed. A network call per prompt would be a bug.
- perfopt 4y agoOK. I noticed that the images are not accurate when I give my own descriptions. Not sure if this is a limitation of Stable Diffusion. For example, for the text "cat and mouse samurai fight in a forest, watched by a porcupine" I got a cat and a mouse (with a cat's face and tail!!) in a forest sort of fighting. But no - porcupine Thank you for creating this.
- bryced 4y agoyes stable diffusion is not great about handling multiple ideas. New image models coming out soon though.
- perfopt 4y agoI keep seeing this even when the prompt is unchanged Downloading https://huggingface.co/runwayml/stable-diffusion-v1-5/resolve/889b629140e71758e1e0006e355c331a5744b4bf/v1-5-pruned-emaonly.ckpt https://huggingface.co/runwayml/stable-diffusion-v1-5/resolv... from huggingface Loading model /home/hrishi/.cache/huggingface/hub/models--runwayml--stable-diffusion-v1-5/snapshots/889b629140e71758e1e0006e355c331a5744b4bf/v1-5-pruned-emaonly.ckpt onto cuda backend... followed by a download
- bryced 4y agoThat is strange. I'm not sure what would cause that unless it was running in some ephemeral environment. What OS? Can you open a github issue with a screenshot?
- GordonS 4y agoWhat are the most affordable GPUs that will run this? (it said it needs CUDA, min 11GB VRAM, so I guess my relatively puny 4GB 570RX isn't going to cut it!)
- bryced 4y agoI'm running on a 2080 TI and an edit runs in 2 seconds. On my Apple M1 Max 32Gb edits take about 60 seconds.
- mritchie712 4y agoIf this was all packaged into a desktop app (e.g. Tauri or electron) how big would the app be? I'd imagine you could get it down to < 500MB (even if you packaged miniconda with it).
- coder543 4y ago> I'd imagine you could get it down to < 500MB (even if you packaged miniconda with it). I don't know where that imagined number came from. This tool appears to be using Stable Diffusion, and the base Stable Diffusion model is 4 or 5 gigabytes by itself. I think there are some other models that are necessary to use the base Stable Diffusion model, and while they are smaller, they still add to the total size.
- ColonelPhantom 4y agoThe cheapest NVidia GPU with 11+GB VRAM is probably the 2060 12GB, although the 3060 12GB would be a better choice. The setup.py file seems to indicate that PyTorch is used, which I think can also run on AMD GPUs, provided you are on Linux.
- pifm_guy 4y agoI really want these ML libraries to get smarter with use of VRAM. Just cos I don't have enough VRAM shouldn't stop me computing the answer. It should just transfer in the necessary data from system ram as needed. Sure, it'll be slower, but I'd prefer to wait 1 minute for my answer rather than getting an error. And if I don't have enough system RAM, bring it in from a file on disk. Tha majority of the ram is consumed by big weights matrices, so the framework knows exactly which bits of data are needed when and in what order, so should be able to do efficient streaming data transfers to have all the data in the right place at the right time. It would be far more efficient than 'swap files' that don't know ahead of time what data will be needed so impact performance severely.
- bryced 4y agoHere is a colab you can try it in. It crashed for me the first time but worked the second time. https://colab.research.google.com/drive/1rOvQNs0Cmn_yU1bKWjCOHzGVDgZkaTtO?usp=sharing https://colab.research.google.com/drive/1rOvQNs0Cmn_yU1bKWjC...
- Tenoke 4y agoHuh, I'm trying it now and the results seem so weak compared to any other model I've seen since dall-e.
- la64710 4y agoHmm that’s true for me too. Not sure if it is due to resource constraint. I had a picture of a car in indoor parking lot with walls and pillars. When I prompted “Color the car blue” the whole image was drenched in a tint of blue. Similarly when prompted “make a hummingbird hover” … the hummingbirds were a patch of shiny colors with an shape that sort of looked like an hummingbird but not like a real one.
- menzoic 4y agoTry "turn the car blue"
- bryced 4y agoDoes dalle do prompt based photo edits now? But yeah sometimes it doesn't follow directions well. I haven't noticed a pattern yet for why that is.
- Damirakyan 4y agoHow would I upgrade to 2.1 if running locally?
- bryced 4y agoIf you're wanting to use Stable Diffusion 2.1 with imaginairy you just specify the model with `--model SD-2.1` when running the `aimg imagine` command.
- TeMPOraL 4y ago> Here are some examples of transformations it can make: Golden gate bridge: I'm on mobile so can't try this myself now. Can it add a Klingon bird of prey flying under the Golden Gate Bridge, and will "add a Klingon bird of prey flying under the Golden Gate Bridge" prompt/command be enough?
- wongarsu 4y agoNo. At least not with the Stable Diffusion 1.5 checkpoint used in the colab notebook. It seems to only have a very vague idea of what a Klingon bird of prey is. The best I could get in ~30 images was [1], and that's with slight prompt tweaks and a negative prompt to discourage falcons and eagles. 1: https://i.imgur.com/gDj2Kn4.png https://i.imgur.com/gDj2Kn4.png
- deleted 4y ago[deleted]
- Daub 4y agoThe language of high-level art-direction can be way more complex than one might assume. I wonder how this model might cope with the following: ‘Decrease high-frequency features of background.’ ‘Increase intra-contrast of middle ground to foreground.’ ‘Increase global saturation contrast.’ ‘Increase hue spread of greens.’
- CyanBird 4y agoThey behave quite poorly, because the keywords used by the models are layman language not technical art or color correction/color grading-speak Hopefully in a couple of years when things have matured more there will be more models capable of handling said requests The most precise models are actually anime models because the users have got high standards for telling the machine what they expect of it and the databases are quite well annotated (booru tags)
- Daub 4y agoAround how many samples are required for an effective training set?
- duringwork12 4y agoLeonardo.ai can probably make a model that handles one of those prompts well with 40 images. To handle them all you would need a larger sample.
- wincy 4y agoWhen I was training Dreambooth on images of myself, then trying to get tags out of the images it generated of me to write better prompts, I clicked “booru tags” on automatic1111 not knowing what it was. It thought I was some sort of Yaoi manga and generated lots of tags that made me both uncomfortable and confused.
- pfd1986 4y agoSuper nice. Would this work if I have my own version of fine-tuned SD? Also, curious how / whether this is different from img2img released by SD. Thanks!
- bryced 4y agoThis is itself it's own finetuned version of SD so now it won't work with alternative versions. img2img works by just running normal stable diffusion img2img on a noised starting image. As such it destroys information at all parts of the image equally. This new model uses attention mechanisms to decide which parts of the image are important to modify. It can leave parts of the image untouched while making drastic changes to other parts.
- xwdv 4y agoHow about “fix the hands”?
- social_quotient 4y agoThe hands issue is going to be an awesome story for all of us in 10-20 years. The younger generation just won’t fathom how hard it was to get proper hands. I wonder what a parallel comparable now would be? Something the slightly technical general public just can’t wrap their head around why it was complicated “back then”. Maybe todays example is a smart voice assistant like Alexa.
- xwdv 4y agoOr maybe it will never be fixed, and in the future when they are trying to determine if someone is a human or an artificial replica, they will simply ask them to draw a set of human hands as a test.
- mensetmanusman 4y agoSomeone needs to redo the blade runner scene with SD with the hands question :)
- AlecSchueler 4y agoMost humans would struggle to draw human hands as well.
- lgas 4y agoWell, this already uses this default negative prompt https://github.com/brycedrennan/imaginAIry/blob/master/imaginairy/config.py#L6 https://github.com/brycedrennan/imaginAIry/blob/master/imagi... so it may fix the hands automatically.
- testtwttttt 4y agoa
- deleted 4y ago[deleted]
- testtwttttt 4y agohow about telling cars where to go ?
- 88stacks 4y agoWow, it's really impressive to see how advanced AI image generators have become! The ability to create stable diffusion images with a "just works" approach on multiple operating systems is a huge step forward in this technology. We've deployed similar tech and APIs for our customers and are contemplating using this library as part of our pipeline for https://88stacks.com https://88stacks.com
- TekMol 4y agoDark patterns are frowned upon here on HN. Letting the user upload dozens of images and only after that telling them they need an account. Not good.
- 88stacks 4y agoits not a dark pattern, what would happen to the images after uploading?
- dspillett 4y ago> its not a dark pattern You are entitled to your opinion, no matter how wrong many of us think it is :) > what would happen to the images after uploading? I have no idea, as there is no such information given going as far as the upload form, nor in the FAQ. This is information you should provide. Though that isn't the key problem IMO. For someone who backs out because of the sign-up requirement, you've wasted their time (and the service now has their images with no obvious pre-agreed policy covering re-use or other licensing issues).
- sschueller 4y agoI am not a fan of software such as this putting in an arbitrarily "safety" feature which can only be disabled via undocumented environment variable. At least make it a flag documented for people who don't have an issue with nudity. There isn't even an indication that there is a "safety" issue, you just get a blank image and are wondering if your GPU/model or install is corrupted. This isn't running on a website that is open to everyone or can be easily run by a novice. Anyone capable of installing and running this is also able to read code and remove such a feature. There is no reason to hide this nor to not document it. Also the amount of nudity you get is also highly dependent on which model you use.
- sandworm101 4y agoNudity isnt really the core issue. It is about illegal content. Nudity is the over-protective, over-inclusive bandaid solution to prevent this thing from being used to generate the very illegal material that will trigger authorities.
- sschueller 4y agoThen instead of just presenting me with a blank image tell me why it's blank. Or add the word "nude" to the default negative filter which by default doesn't have that in it. The default is: negative-prompt:"Ugly, duplication, duplicates, mutilation, deformed, mutilated, mutation, twisted body, disfigured, bad anatomy, out of frame, extra fingers, mutated hands, poorly drawn hands, extra limbs, malformed limbs, missing arms, extra arms, missing legs, extra legs, mutated hands, extra hands, fused fingers, missing fingers, extra fingers, long neck, small head, closed eyes, rolling eyes, weird eyes, smudged face, blurred face, poorly drawn face, mutation, mutilation, cloned face, strange mouth, grainy, blurred, blurry, writing, calligraphy, signature, text, watermark, bad art,"
- sandworm101 4y agoSure, if this was a commercial product i would also scream about feedback. But as this is a free side project im not going to criticize decisions on what they think they need to do to avoid unnecesary drama. I credit them for making the bypass relatively simple.
- weakwire 4y agoEnchance!
- distantsounds 4y agoIf only Stable Diffusion wasn't already populated with a host of copyrighted images already. Make your own art, dammit. This is the equivalent running some Photoshop filters through someone else's work.
- social_quotient 4y agoSlightly off topic. I’ve been looking for an easier way to replace the text in these ai generated images. I found Facebook is working on it with their TextStyleBrush - https://ai.facebook.com/blog/ai-can-now-emulate-text-style-in-images-in-one-shot-using-just-a-single-word/ https://ai.facebook.com/blog/ai-can-now-emulate-text-style-i... but have been unable to find something released or usable yet. Anyone aware of other efforts?
- johndough 4y agoThe authors of TextStyleBrush cite SRNet, which is available at https://github.com/youdao-ai/SRNet https://github.com/youdao-ai/SRNet but probably has worse quality. I don't know of others, but I have not looked very hard either.
- sfpotter 4y agoThese look awful! They are very displeasing aesthetically. They look like they were done by someone with absolutely no artistic ability. Clearly there is some technical interest here, but I just felt the need to point out the elephant in the room. They are very ugly.
- aflag 4y agoI'm not an artist, but they look fine to me. I am not the kind of person who spends hours in a Gallery admiring the nuances of paintings or photos. However, at the level of detail I usually admire these things, the clown one looked interesting and the Monalisa one was funny. The strawberry seemed a bit weird, but I don't think I'd care for it even if it was perfect anyway. The wintery landscape I thought was pretty good and the red dog it delivered what was asked. Not sure how it could be much different than that.
- sfpotter 4y ago"I'm not qualified to have a nuanced opinion about this, but let me confidently tell you what I think..."
- sh4rks 4y agoI love the cognitive dissonance between "you're just stealing people's art and modifying them slightly!" versus "AI art sucks and has no artistic value"
- theusus 4y agoTwo things 1. It actually makes me insecure. 2. Don't we already have apps that do such things? Yes, they were more specialized, but it's the same thing as Prisma app.
- tomrod 4y agoThis is a lot of fun! And they aren't kidding that on a CPU backend it is slooooow :)
- sandworm101 4y agoFireworks. These AI tools seem very good at replacing textures, less so about inserting objects. They can all "add fireworks" to a picture. They know what fireworks look like and diligently insert them into "sky" part of pictures. But they don't know that fireworks are large objects far away rather than small objects up close (see the Father Ted bit on that one). So they add tiny fireworks into pictures that don't have a far away portion (portraits) or above distant mountain ridges as if they were stars. Also trees. The AI doesn't know how big trees are and so inserts monster trees under the Golden Gate bridge and tiny bonsais into portraits. Adding objects into complex images is totally hit and miss.
- taberiand 4y agoPerhaps stereoscopic video should be part of the training data?
- YurgenJurgensen 4y agoHuman stereoscopy is only good out to a few meters (and presumably people aren't going out with giant WWII stereoscopic rangefinders to generate training data). So it wouldn't help them for things like fireworks or trees.
- treeman79 4y agoFound out I couldn’t see in stereo. Got prism glasses. Completely blew my mind to seen ”depth” for the 1st time. Had no idea I couldn’t. Never had any trouble. Apparently without prism glasses my vision just switches from one eye to the next every 30 seconds. Completely seamlessly.
- dr_dshiv 4y agoWow! Any brand suggestions? And, are these the same prism glasses as those that let you watch tv laying in bed?
- 4y ago
- fassssst 4y agoRelated: https://www.reddit.com/r/StableDiffusion/comments/10hv160/image_editing_with_just_text_prompt_new/ https://www.reddit.com/r/StableDiffusion/comments/10hv160/im...
- nmstoker 4y agoLooks really interesting, although my immediate thought with "-fix-faces" is how long before someone manages to do something inappropriate and whip up a storm about this.
- dandigangi 4y agoThis is really cool. Haven't seen something like this yet. Going to be very interesting when you start to see E2E generation => animation/video/static => post editing => repeat. Have this feeling that movie studios are going to look into this kind of stuff. We went from real to CGI and this could take it to new levels in cost savings or possibilities.
- dandigangi 4y agoPlayed around for a bit. Definitely a cool tool. Wish I had an M1 though. Taking me quite a bit to generate and fans running at full blast. Haha
- lou_alcala 4y agoWow this is cool I think I am going to make a site so people can use this
- cbeach 4y agoI see it's able to generate politician faces. I recall this wasn't possible on DALL·E 2 due to safety restrictions. I run a friendly caption contest https://caption.me https://caption.me so imaginAIry is going to be absolute gold for generating funny and topical content. Thank you @bryced!
- karim79 4y agoI've been toying with SD for a while, and I do want to make a nice and clean business out of it. It's more of a side-projecty thing so to speak. Our "cluster" is running on a ASUS ROG 2080Ti external GPU in the razer core-x housing, and that actually works just fine in my flat. We went through several iterations of how this could work at scale. The initial premise was basically the google homepage, but for images. That's when we realised that scaling this to serve the planet was going go be a hell of a lot more work. But not really, conceptualising the concurrent compute requirements as well as the ever-changing landscape and pace of innovation in this absolutely necessary. The quick fix is to use a message queue (we're using Bull) and make everything asynchronous. So essentially, we solved the scaling factor using just one GPU. You'll get your requested image, but it's in a queue, we'll let you know when it's done. With that compute model in place, we can just add more GPUs, and tickets will take less time to serve if the scale engineering is proper. I'm no expert on GPU/Machine learning/GAN stuff but Stable Diffusion actually prompted me to imagine how to build and scale such a service, and I did so. It is not live yet, but when it does become so the name reserved is dreamcreator dot ai, and I can't say when it will be animated. Hopefully this year.
- throwaway675309 4y agoA ticketing scheduler system is how 99% of systems that require long running CPU/GPU intensive that cannot be run in parallel are implemented. It's how I built up my stable diffusion discord bot which is backed by a single RTX 2060. I'm glad you have this working but I wouldn't exactly call this "solving the scaling problem", you're just running it in a blocking "serial fashion". With enough concurrent users it could still take somebody until the heat death of the universe for their image to finally be generated.
- karim79 4y agoAgreed. The scaling model/queueing system was implemented as a POC which can be scaled by plugging in more cards/hosts. I hope we can find the time to animate this soon enough.
- DrScientist 4y ago
- anigbrowl 4y agoA CUDA supported graphics card with >= 11gb VRAM (and CUDA installed) or an M1 processor. /Sighs in Intel iMac Has anyone managed to get an eGPU running under MacOS? I guess I could use Colab but I like the feeling of running things locally.
- lightbulbish 4y agoThis is cool! Makes me want to pull the trigger on an M2
- fatih-erikli 4y agoGarbage.
- mstade 4y agoDoes anyone know of any tool like this for UI design? I'd love something that'd help creatively impaired people like myself communicate more visually.
- goffi 4y agoWow that's really impressive (I've seen similar things in research papers for a while now, but having it usable so easily and generic is great). A few questions: - would it be possible to use this tool to make automatic mask for editing in something like GIMP (for instance, if I want to automatically mask the hair)? - would it be possible to have a REPL or something else to make several prompt on the same image? Loading the model takes time, and it would be great to be able to just do it once. - how about a small GUI or webui to have the preview immediately? Maybe it's not the goal of this project and using `instruct-pix2pix` directly with its webui is more appropriate? Thanks for the work (including upstream people for the research paper and pix2pix), and for sharing.
- bryced 4y ago> would it be possible to use this tool to make automatic mask for editing in something like GIMP probably but GIMP plugins are not something I've looked into > REPL already done. just type `aimg` and you're good to go > GUI GUIs add a lot of complexity. Can your file manager do thumbnails and quick previews?
- sseagull 4y ago> GUIs add a lot of complexity. Can your file manager do thumbnails and quick previews? Somewhat OT, but I find this really funny. It says a lot about the difficulty of using various ecosystems and where communities spend time polishing things. "Yeah, I made something that takes natural language and can do things like change seasons in an image. But a GUI? That's complicated!" It's not a criticism of you, but the different ecosystems and what programmers like to focus on nowadays.
- bryced 4y agoFair but I'd point out I also didn't make the algorithm that changes photos. I'm wrapping a bunch of algorithms that other people made in a way that makes them easy to use. It's not just that GUI's are hard, it's that the "customer" base will inevitably be much less technical and I'd receive a lot more difficult to resolve bug reports. So no-gui is also a way of staying focused on more interesting parts of the project.
- petrusnonius 4y agoAre you telling me I can finally ENHANCE!? Great stuff man, thanks!