18 ms·
I used Stable Diffusion and Dreambooth to create an art portrait of my dog
- simonw 3y agoI love how much work went into this. There's a great deal of pushback against AI art from the wider online art community at the moment, a lot of which is motivated by a sense of unfairness: if you're not going to put in the time and effort, why do you deserve to create such high equality imagery? (I do not share this opinion myself, but it's something I've seen a lot) This is another great counter-example showing how much work it takes to get the best, deliberate results out of these tools.
- minimaxir 3y agoUnfortunately it's become a meme among AI art haters that AI art is "just inputing text into a text box" despite the fact that is far from the truth, particularly if you want to get specific results as this blog post demonstrates. Some modern AI art workflows often require more effort than actually illustrating using conventional media. And this blog post doesn't even get into ControlNet.
- tester457 3y agoIt's a meme because 99% of the ai art creators don't go that deep, they only prompt. Even if they did have a more complex workflow most of them are still based on copyrighted training data, so there will be many lawsuits.
- capableweb 3y ago> Some modern AI art workflows often require more effort than actually illustrating using conventional media. And this blog post doesn't even get into ControlNet. Indeed. Another criticism that I can definitely somewhat see the idea behind, is that the barrier to entry is very different from for example drawing. To draw, you need a pen and a paper, and you can basically start. To start with Stable Diffusion et al, you need either A) paid access to a service, B) money to purchase moderately powerful hardware or C) money to rent moderately powerful hardware. One way or another, if you want to practice AI generated art, you need more money than what a pen and paper cost.
- realusername 3y agoStable Diffusion doesn't really need powerful hardware, any graphic card will do, it will just be a bit longer. There's even ports on smartphones nowadays.
- simonw 3y agoThere are plenty of traditional art mediums that require significant financial outlays to get started: oil painting, ceramics, glass blowing etc. There are plenty of free online tools for using all kinds of AI image generation techniques, and they don't require powerful hardware, just something that can browse websites or run Discord.
- adamm255 3y agoPlus training, lessons and inspiration. And talent. It’s like with dreams. They can be terribly intricate and detailed, but ask me to draw something creative and I’m out.
- wincy 3y agoI've probably spent at least an hour a day working in Stable Diffusion and Automatic1111 since January or so. At this point I'd call it my hobby, as instead of playing video games I'm plowing time into this. And I'm definitely seeing marked improvements in my style and what I'm looking for. I'll often start with a shotgun approach and make 64-128 pictures with a basic prompt of what I'm looking for. Sometimes there's a shape and basic composition that makes me gasp it's so much better than the others. So I'll feed that into img2img or inpainting, iterate on it, tweak settings that I just sort of have an intuitive feel for what turning the knobs does, and while away for an hour or two to make it exactly what I want. There's definitely a "dreamy fantasy art style" I'm a big fan of that would have cost me an absolute fortune to commission even one image a year ago. I can't match what is coming from the top artists on Artstation (I tried a few weeks back which was humbling, some astounding work). But it's good enough that my D&D buddies are entertained and amazed that I went through and made each one of their characters. Our DM, being someone who has released creative works, was reticent and less gung-ho on AI for awhile until he decided to start playing with the AI tools (Midjourney in his case) for a new custom campaign he's running. He's suddenly able to develop novel NPC tokens for every important character we meet, the monsters are high resolution and the convenience of using Midjourney in Discord (which he already uses to coordinate our online D&D games) has been a huge boon and enhancement to how much fun and immersive our games are. A year ago this would have cost literally tens of thousands of dollars. He's a published fantasy author so prompting aka describing a scene comes naturally to him. It's been a lot of fun seeing what he's coming up with. I'm really loving the spark of creativity I've been finding in myself where the turnaround on the old tools was too long for me to not get frustrated and give up, and to see it amongst my friends, even the ones who were initially skeptical.
- squidsoup 3y agoOnly if you exclude the countless hours an illustrator has spent developing their craft.
- yieldcrv 3y agobeing sympathetic to that requires pretending that the user would have ever commissioned an artist for that idea at all. both the transaction and the idea would have simply never happened. it was never valuable enough or important enough to commission a human, hope you got the correct human, wait week after week for revision after revision. people that want to hone a niche discipline for themselves still can do that. just be honest about doing it for yourself.
- libraryatnight 3y agoUsing AI as a tool to create art takes nothing away from anyone who spent time learning a skill or craft that they use in their own pursuit of expression. People will be arguing about whether or not art made with AI is art, and artists will just be using it or not. I remember an interview about electronic music where Bjork addressed concerns that if you use a computer to make music, it has no soul, and she said if the person using the machine to make the music puts soul into it, it will have a soul. I remember David Bowie in the mid 90s saying if he was young in that decade he might not have been a musician, because in the 60s being a musician seemed subversive and at the time of the interview the internet was carrying the flag of subversion. Anyway, it's interesting to watch these conversations. I'd never claim to know what art is or try to tell someone, but it seems to me that already because of the controversy artists are drawn to AI and further exciting the conversation. Commercial artists seem the most threatened; animators, designers, etc. I understand why, but I don't think arguing that AI isn't "art" is going to help their cause any more than protesting digital painting wasn't art, electronic music wasn't art, and much earlier that photography wasn't art. All the time these conversations are happening, the art's getting made and we're barreling towards the next 'not art' movement.
- jachee 3y ago> Using AI as a tool to create art takes nothing away from anyone who spent time learning a skill or craft that they use in their own pursuit of expression. Except all those artists’ art being used without their consent to train these models that subvert the exclusivity of their style, or obviate their work altogether. It ingests their literal effort and eliminates other humans’ need to put forth a similar level of effort. I’m all for cool new tools, but this is very much like the invention of digital sampling, and models should be required to “clear” all the works that they “sample” for training.
- dorkwood 3y ago> Some modern AI art workflows often require more effort than actually illustrating using conventional media. Then why don’t they illustrate it instead, and save themselves some time?
- efreak 3y agoWhy don't you buy your cake and cookies at a bakery instead of making them yourself at home?
- dorkwood 3y agoI’m not sure I understand your argument. If you’re suggesting that illustrating by hand requires more effort than automatically generating an image with AI, then we agree with each other.
- mdp2021 3y agoThe shruggingface submission is very interesting and very instructive. Nonetheless, it would be odd and a weak argument to point criticism towards not spending adequate «time and effort» (as if it made sense to renounce tools and work through unnecessary fatigue and wasting time). More proper criticism could be in the direction of "you can produce pleasing graphics but you may not know what you are doing". This said, I'd say that Stable Diffusion is a milestone of a tool, incredible to have (though difficult to control). I'd also say that the results of the latest Midjourney (though quite resistant to control) are at "speechless" level. (Noting in case some had not yet checked.)
- Paul-Craft 3y ago> More proper criticism could be in the direction of "you can produce pleasing graphics but you may not know what you are doing". I don't get this. If one "can produce pleasing graphics," how does that not equal knowing what they're doing? I only see this as being true in the sense of "Sure, you can get places quickly in a car, but you don't really know how it works."
- mdp2021 3y ago> how does that not equal knowing what they're doing The goal may not be to produce something pleasant. The artist will want some degree of artistic value; the communicator will want a high degree of effectiveness etc. The professional will implicitly decide a large number of details, in a supposedly consistent idea of the full aim. The non professional armed with some generative AI tool may on the contrary leave a lot to randomness - and obtain a "pleasant" result, but without real involvement, without being the real author nor, largely, the actual director.
- bawolff 3y agoThat seems untrue. In this case, the author set out with a specific goal, then tried to do it, and then succeded. What are you suggesting, that the author lied in the blog post and actually worked backwards, post hoc? Seems incredibly unlikely based on the details they wrote.
- asddubs 3y agomost of the criticism I've seen is that it's all trained on uncompensated stolen artwork. Much like how copilot is trained on GPL code, disregarding its license terms.
- simonw 3y agoThe trained on stolen artwork critique is reasonable - I helped with one of the first big investigations into how that training data worked when Stable Diffusion first came out: https://simonwillison.net/2022/Sep/5/laion-aesthetics-weeknotes/ https://simonwillison.net/2022/Sep/5/laion-aesthetics-weekno... It's interesting to ask people who are concerned about the training data what they think of Adobe Firefly, which is strictly trained on correctly licensed data. I'm under the impression that DALL-E itself used licensed data as well. I find some people are comfortable with that, but others will switch to different concerns - which indicates to me that they're actually more offended by the idea of AI-generated art than the specific implementation details of how it was trained.
- bugglebeetle 3y agoI think the more correct argument is that Stable Diffusion effectively did a Napster to force artists into shit licensing deals with large players who can handle the rights management. It’s unlikely that artists would’ve ever agreed to them otherwise, but since the alternative now is to have your work duplicated by a pirate model or legally gray service, what are you going to do? This seems borne out by the fact that Stability AI themselves are now retreating behind Amazon for protection.
- adamm255 3y agoWhen I did Photography at college, a lot of the work was looking at other works of art. I spent a lot of time in Google Images, diving through books from the Art section and going to galleries. Lots of photo copying was involved! I then did works in the style of what I’d researched. I trained myself on works I didn’t own, and then produced my own. I kind of see the AI training as similar work, just done programmatically vs physically. Certainly a very interesting topic. I can’t get my head around how far we’ve come on this in the last 6-12 months. From pretty awful outputs to works winning Photography awards. And prints of a dog called Queso you’d have paid a lot of money to an illustrator for.
- brucethemoose2 3y agoTBH it would be much easier with more streamlined tooling, especially if doing it locally with lora/lycoris. Its kinda like using ffmpeg for vapoursynth for video editing instead of a video editing GUI. That being said the training parameter/data tuning is definitely an art, as is the prompting.
- quadcore 3y agoa lot of which is motivated by a sense of unfairness Say you generate a picture with midjourney - who is/are the closest artist(s) you can find for that picture? Not the AI, not the prompter, so the closest artists you can find for that picture are the ones who made the pictures in the training set. So generating a picture is outright copyright infringement. Nothing to do with unfairness in the sense of "artists get out compete". Artists dont get out compete - they are stolen.
- ModernMech 3y agoTypical Midjourney workflow involves constantly reprompting and fine tuning based on examples and input images. When you arrive at a given image in Midjourney, it’s often impossible to recreate it even with the same seed. You’ll need the input image as well, and the input image is often the result of a long creative process. Why is it you discount the creative input of the user? Are they not doing work by guiding the agent? Don’t their choices of prompt, input image, and the refinement of subsequent generated images represent a creative process?
- quadcore 3y agoI agree with you on the technicality - if we say the promter is an artist, then the picture belongs to him.
- quadcore 3y agoFrom what I read on the internet, people assume AI generated art is a difficult question legaly speaking. Some literally assume artists complain only because there are out competed. I disagree - I think that AI generative art is an easy case of copyright infrigement and an easy win for a bunch of good lawyers. That's because you can't find an artist for a generated picture other than the ones in the training set. If you can't find a new artist, then the picture belongs to the old ones, so to speak. I really dont see what's difficult with that case. I think the internet assume a bit to quickly it's a difficult question and a grey area when maybe it just isnt. It's noteworthy that Adobe did things differently than the others and the way they did things goes in the direction im describing here. Maybe it's just confirmation bias.
- circuit10 3y agoIt’s not as simple as that though because the algorithm does learn by itself and mostly just uses the training data to score itself against, it doesn’t directly copy it as some people seem to think. It can end up learning to copy things if it sees them enough times though “you can't find an artist for a generated picture other than the ones in the training set. If you can't find a new artist, then the picture belongs to the old ones, so to speak” I don’t think that’s valid on its own as a way to completely discount considering how directly it’s using the data. As an extreme example, what if I averaged all the colours in the training data together and used the resulting colour as the seed for some randomly generated fractal or something? You could apply the same arguments - there is no artist except the original ones in the training set - and yet I don’t think any reasonable person would say that the result obviously belongs to every single copyright owner from the training set
- ModernMech 3y agoBut this person’s dog isn’t in the training set, so why should some artist be credited for a picture they never drew? Not a single person has drawn his dog before, now there is a drawing of his dog, and you want to credit someone who had no input to the creative process here?
- quadcore 3y ago
- basisword 3y ago> if you're not going to put in the time and effort, why do you deserve to create such high equality imagery? This isn’t high quality imagery. Don’t get me wrong, the tech is cool and I love the work that’s went into making this picture. But this isn’t something I would ever hang on my wall. There’s probably a market for it, but I get the strong impression it’s the “live, laugh, love” market. The people that buy pictures for their wall in the supermarket. The kind of people who pay individual artists money to paint bespoke images of their pet are not going to frame AI art. I don’t think the artists need to worry.
- yellow_postit 3y agoI would expect it’s only a matter of time till those “traditional” artists also adopt these tools into their workflows. Similar to the initial pushback against the “digital darkroom” which is now the mainstay of photography. In-ai-aided art, like manually developed film, will trend towards a niche.
- theaiquestion 3y ago> This isn’t high quality imagery. Don’t get me wrong, the tech is cool and I love the work that’s went into making this picture. But this isn’t something I would ever hang on my wall. Well yeah but that doesn't change the OP commenter's point that it takes a lot of work to get high quality art still. > I don’t think the artists need to worry. I disagree here but only on the basis of what type of art it is. Stock art/photography, and a lot of media designwork is likely at risk because we can now create "good enough" art at the click of a button for almost no cost. I agree that the "hang on the wall level good" artists aren't at risk just yet, but between the more filler-art and the uh Well "anime/furry" commissioners are definitely at risk right now for anything except the highest quality artists, and there is a MASSIVE community behind this - in fact they have done a lot of the innovation for StableDiffusion including optimizations/A1111 webui, and have trained many custom models for their art, already had pretagged datasets of 10k's of images....
- syntheweave 3y agoSimpler cartooning styles like classical Disney/Warner are actually some of the weakest ones in AI models right now. Prompted diffusion models are poorly suited to the task and tend to capture excess detail in the contour, because they don't have a mechanism for calculating out the geometry of a constructed form, so they don't arrive at the same clarity and simplicity. What the model can do best is to convert a basic contour sketch into an elaborate rendering, which means that it's the top end of the market that was doing those renderings that's most at risk.
- davely 3y agoI love the detailed workflow that OP posted. Dogs seem to be particularly good subject material for this. I turned my dog into a robot awhile back using the img2img feature of Stable Diffusion and the results were pretty amazing![1] [1] https://twitter.com/davely/status/1583233180177297408 https://twitter.com/davely/status/1583233180177297408
- madeofpalk 3y ago> a lot of which is motivated by a sense of unfairness This is not something I've seen once in any sort of criticism of "AI art", and elsewhere in the internet I'm largely in a anti-ai-art bubble. Most legitimate pushback I've seen has been more on the non-consensual training of models. Many artists don't want their work to be sucked up into the "AI Borg Model" and then regurgitated by someone else, removing the artists consent, credit, and compensation.
- regularfry 3y agoI absolutely have seen it. A lot. It's dressed up as Luddism, more often expressed as "you shouldn't be able to have those results because I spent years honing my craft" which may or may not be followed by "...and if we allow this, those years were wasted and I'm out of a job, along with millions of others".
- TeMPOraL 3y ago> It's dressed up as Luddism > "...and if we allow this, those years were wasted and I'm out of a job, along with millions of others" Given that the quoted part is highly likely to happen, and not just in art, I hope we're no longer considering "luddism" to be a pejorative.
- bawolff 3y agoLuddites were a real group, who really did lose their jobs to technical progress. So it seems fair. Technical progress does really require adaption sometimes. We don't criticize luddites because we think they were wrong about change being a real thing.
- odessacubbage 3y ago[flagged]
- jhbadger 3y ago
- Auracle 3y agoI've done so much with a fine-tuned model of my dog. I previously made coloring pages for my daughter of our dog as an astronaut, wild west sheriff, etc. They're the first pages she ever "colored," which was pretty special for us. Currently I'm working on making her into every type of Pokemon, just for fun.
- mdp2021 3y agoUsing which tools, specifically?
- Auracle 3y agoStable Diffusion, generically. StableTuner to fine tune the model - I can't recall the name of the model I trained on top of, but it was one of the top "broad" 1.5 based models on Civitai. Automatic1111 to do the actual generating. I used an anime line art LoRA (at a low weight) along with an offset noise LoRA for the coloring book pages as otherwise SD makes images be perfectly exposed. For something like that you obviously want a lot more white than black. EveryDream2 would be another good tuning solution. Unfortunately that end of things is far from easy. There are a lot of parameters to change and it's all a bit of a mess. I had an almost impossible time doing it with pictures of my niece, my wife is hit or miss, her sister worked really well for some reason, and our dog was also pretty easy.
- Auracle 3y agoI uploaded a couple of the Pokemon generations really quick as examples. I still need to go through and do quick fixes for double tails (the tails on Pokemon are not where they are on regular animals, apparently), watermarks, etc. and do a quick Img2Img on them. https://imgur.com/a/11OxoSA https://imgur.com/a/11OxoSA
- minimaxir 3y agoFor generating Pokemon, I recommend using this model along with a textual inversion of your pet: https://huggingface.co/lambdalabs/sd-pokemon-diffusers https://huggingface.co/lambdalabs/sd-pokemon-diffusers
- sinman 3y agoI did something loosely related. As a present for my girlfriend's birthday, I made her a "90s website" with AI portraits of her dog: https://simoninman.github.io/ https://simoninman.github.io/ It wasn't actually particularly hard - I used a Colab notebook on the free tier to fine-tune the model, and even got chatGPT to write some of the prompts.
- Auracle 3y agoIn my (limited) experience, dogs seem to be easier than people for fine-tuning - especially if your end result is going to be artsy. Faces of people you know well being off in slight ways really throws you off, but with dogs there's a bit more leeway.
- jakedahn 3y agohah, these are pretty cool! Well done!
- itronitron 3y ago[flagged]
- spikej 3y agoThey like it. And it was a good excuse to work with new tech. Why poo poo on it?
- itronitron 3y agoI like my comment, and it was a good excuse to work with new tech, why poo poo on my comment?
- mdp2021 3y agoBecause your comment was pretty objectively inappropriate and improductive - gratuitous. Or did you mean something productive that we should have guessed?
- Fricken 3y agoLeaving poo poo on things is a popular passtime for many dog people.
- mdp2021 3y ago'is a popular passtime for many [] people' Fixed That For You. Interestingly for the ethologist, they have habitats: for example, the bottom comments in the stacks in YouTube...
- mdp2021 3y agoGood, so in order to produce good AI-aided graphics the producers will have to become critics, arts experts, with the important side effect of personal elevation and the collective gain of society. "Wins" on all sides. Update: three minutes later, it seems that somebody did not get the irony.
- beezlewax 3y ago
- wincy 3y agoHe mentions the Colab for Dreambooth, that only takes ten minutes or so to train using an A100 (the premium GPU) and you can have it turn off after it finishes, and saves to Google Drive. Super easy.
- jakedahn 3y agoYeah! Here's the colab notebook, in case anyone is interested: https://github.com/TheLastBen/fast-stable-diffusion https://github.com/TheLastBen/fast-stable-diffusion I've trained a few smaller models using their Dreambooth notebook, but I think for 4000 training steps, an A100 will usually take 30-40min. I believe replicate also uses A100s for their dreambooth training jobs.
- wincy 3y agoAh I see, you're right 40 minutes sounds about right for that amount of training. Curious why the decision to train 40 images? I've used 15 for two separate subjects in Dreambooth with excellent results. I'm no expert, experimenting the same way as you, but haven't trained on more than 15-20 images per subject. I've found the most important part is spending a good amount of time getting the prompts, although I'm not sure if having the person in an environment embodied and describing the objects around them helps give the model a "sense of scale"? For example if I just train "wincy" in the fast Dreambooth "wincy" will be the only token it'll know, with no other info in the prompts, it didn't know what in the image was "wincy" (me). I accidentally did this on training my wife (no prompts at all) and she got really mad at me at how ugly the results were (you made me ugly! haha) Have you tried it with and without your dog in an environment, then describing the environment your dog is in for the training data?
- MasterScrat 3y agoFYI we're building a service to make this process even simpler and faster: dreamlook.ai Upload your pictures, we train the model in a few minutes, then you can download your trained checkpoint. $1/model, first one for free. For app builders, we provide a solid API that scales to 1000s of runs per day without breaking a sweat.
- asadlionpk 3y agoIf anyone wants to try Dreambooth online, I made a free website for this: https://trainengine.ai https://trainengine.ai
- cinntaile 3y agoIt's unfortunate a lot of the nice artsy detail disappeared when he had to recreate part of the head, but I guess that is inevitable. Great work and interesting writeup.
- EGreg 3y agoWhat are the tools we can run on a Linux machine? EDIT: four downvotes and zero answers how to run it on a Linux machine…
- cogitoergofutuo 3y agoThe only piece of software mentioned in the article that doesn’t run on Linux is Draw Things.
- minimaxir 3y agoYou were likely downvoted because you asked how to use it for NFTs, which you just edited it out.
- EGreg 3y agoI don’t see why that is relevant. Why is using it for NFTs worthy of a downvote?
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- liuliu 3y agoThere might be a few things Draw Things missing from this article: no mask blur, not selecting the inpainting model for inpainting work. Tomorrow's release should contain both mask blur and inpainting ControlNet, which might help these use cases.
- jakedahn 3y agoYeah, it was likely just user error. I actually really love Draw Things, because I can run it locally on my mac and quickly experiment without having to sling HTTP requests or spin up GPUs. I did the actual work back on March 11th, so I was likely on an older build; but I was seeing issues where inpainting was just replacing my selection/mask with a white background. I had the inpainting model loaded, but couldn't figure it out. I'm planning to continue playing with Draw Things locally, and exploring the inpainting stuff. For such an iterative process I feel like a local client would make for the best experience.
- liuliu 3y agoThere is no user error but UX issues :) That has been said, you probably used paintbrush rather than the eraser? There would be more help on the Discord server though! https://discord.gg/5gcBeGU58f https://discord.gg/5gcBeGU58f
- bigbillheck 3y agoPersonally I paid a friend $200 to create an art portrait of my dog.
- mkoryak 3y agoNot all of us have friends or 400$
- deleted 3y ago[deleted]
- indigodaddy 3y agoPretty cool stuff. Personally though, not a huge fan of his “the one” choice. Some of the other images in his assortment were much better imo. Each to their own of course though!
- steve_adams_86 3y agoI agree, but I find it pretty cool that they were able to generate and pick from what they wanted. This seems like one of the real strengths of generative AI — people can tune outputs they otherwise couldn’t create (unable to paint, draw, play guitar, etc). People can debate if it’s actually good that people can create art without being artists, but again, I think it’s great that the author had the freedom to create what they had in mind without much outside influence. This has been a goal for computers in general for so long, and it seems like we’re actually arriving with some mediums.
- sammalloy 3y ago> Pretty cool stuff. Personally though, not a huge fan of his “the one” choice. Some of the other images in his assortment were much better imo. Each to their own of course though! Glad to see I’m not alone on this. I think the end result would have turned out much better if the author had simply adhered to the Huichol art palette, which I’m convinced they were aiming for at the beginning. That color scheme works for a reason. https://en.wikipedia.org/wiki/Huichol_art https://en.wikipedia.org/wiki/Huichol_art
- amelius 3y agoBut why pick a dog as an example? Humans are much worse in telling dogs apart than other humans (except perhaps the owner of the particular dog). So for all we know, the AI didn't generate a portrait of this particular dog but instead a generic picture of this breed of dog.
- chipgap98 3y agoBecause you invent a new word when you train dreambooth and teach it that your subject is an example of that word. The fact that the word you've created returns photos similar to subject is a sign that it worked.
- amelius 3y agoI suppose that dreambooth is pretrained on a large dataset that includes many different dogs. My point is that it is difficult to judge (for us) that the returned photos are actually similar to the subject.
- ModernMech 3y agoThe paper shows dogs with very distinctive fur coloring. Particularly the corgi with a white strip between its eyes. I think the paper would be completely fraudulent if this dog were also featured heavily in the training set. So the point is the white stripe corgi isn’t in the set, and with a few examples, the model could then generate brand new images of corgis with a similar fur pattern. Maybe all it can do is fur patterns but it’s a start.
- jakedahn 3y agoMostly because I thought of it more as an art project than a technical accuracy project. However, the honest answer to your question, is because I have a ridiculous amount of photos of my dog on my phone . Getting training data is hard work. But this is totally true, I found that maybe 30% of the images I generated did not look like my dog at all. However the rest do a good job at capturing his eyes and facial expressions that he actually makes. I thought that the chosen image I worked from captured the look of his eyes super well. But yeah, nobody but me would really appreciate that.
- lxe 3y agoThis is a great writeup on some of the nuances and gotchas you have to watch out for when finetuning using dreambooth and the generative creative process in general.
- spaceman_2020 3y agoI would highly recommend using Photoroom's background removal tool. Does a far, far better job than Photoshop.
- throwaway675309 3y agoPixelMator is a highly competitive native Mac app, has an excellent background remover and unlike photoroom/PS, it's a one time purchase.
- cogitoergofutuo 3y agoThis is really interesting. I do wish the author included the cost to train the model from replicate though.
- yieldcrv 3y agoResults at the top of your article/project please
- skor 3y agoI liked the original more than the final version. The vector style drawing was much more futuristic and more interesting. Seems like lots of work went into that and I hope the author enjoyed the process and enjoys the final result.
- joegahona 3y agoI did too, and I even liked the aggressive cropping. Totally subjective though. The final result was beautiful as well, and this was a joy to read.
- birdfood 3y agoI find it interesting that this green/orange colour palette so commonly appears in midjourney images, seemingly regardless of the subject.
- cubefox 3y agoNot just Midourney, also Dall-E 3 often does it: https://www.bing.com/images/create/a-bright-day-on-the-streets-of-a-buzzing-cyberpunk/641cb127043e41fdbc79087f806e4f48?id=pp5XL8a4wfbe1ghc8FP%2fxA%3d%3d&view=detailv2&idpp=genimg&FORM=GCRIDP&mode=overlay https://www.bing.com/images/create/a-bright-day-on-the-stree... https://www.bing.com/images/create/christmas-on-board-a-spaceship2c-dslr-photograph/643b11f856344beda54cf8d145625a36?id=2QI%2f5m5raxPlSmsxYWME%2bQ%3d%3d&view=detailv2&idpp=genimg https://www.bing.com/images/create/christmas-on-board-a-spac... https://www.bing.com/images/create/in-the-streets-of-atlantis2c-photorealistic2c-dslr/6439c8a55de845d8a90fee5b0affae1c?id=ROcvehIdRCtcZpr5bi%2fszQ%3d%3d&view=detailv2&idpp=genimg https://www.bing.com/images/create/in-the-streets-of-atlanti...
- GaggiX 3y ago>I was wrong . it seemed to take the top and bottom-most row of pixels and extend them down from 512px tall to 1344px tall. I mean you cannot outpaint in the img2img tab, load the image in the inpaint tab and possibly use the inpainting model.
- jakedahn 3y agoah-ha! This was probably it
- pjgalbraith 3y agoSome feedback on workflow: - Automatic1111 outpainting works well but you need to enable the outpainting script. I would recommend Outpainting MK2. What the author did was just resize with fill which doesn't do any diffusion on the outpainted sections. - There are much better resizing workflows, at a minumum I would recommend using the "SD Upscale Script". However you can get great results by resizing the image to high-res (4-8k) using lanczos then using inpainting to manually diffuse the image at a much higher resolution with prompt control. In this case "SD Upscale" is fine but the inpaint based upscale works well with complex compositions. - When training I would typically recommend to keep the background. This allows for a more versitile finetuned model. - You can get a lot more control of final output by using ControlNet. This is especially great if you have illustration skills. But it is also great to generate varitions in a different style but keep the composition and details. In this case you could have taken a portrait photo of the subject and used ControlNet to adjust the style (without and finetuning required).
- jakedahn 3y agoThank you for these recommendations! I'll definitely be trying them next time 'round.
- pjgalbraith 3y agoGood luck! I have some workflow videos on Youtube https://youtube.com/pjgalbraith https://youtube.com/pjgalbraith. But I haven't had a chance to show off all the latest techniques yet.
- raincole 3y ago> However you can get great results by resizing the image to high-res (4-8k) using lanczos then using inpainting to manually diffuse the image at a much higher resolution with prompt control. Diffuse an 8k image? Isn't it going to take much, much more VRAM tho?
- SV_BubbleTime 3y agoThat confused me at first too.You aren’t diffusing the 8k image. You are upsampling, then inpainting sections that need it. So if you took your 8K and inpainted a section with 1024x1024 that works well with normal ram usages. In Auto1111, you need to select “inpainted masked area” to do that.
- hartator 3y agoIsn’t disappointing that nothing important is open source these days in AI?
- cogitoergofutuo 3y agoDreambooth implementation https://github.com/JoePenna/Dreambooth-Stable-Diffusion https://github.com/JoePenna/Dreambooth-Stable-Diffusion AUTOMATIC1111 https://github.com/AUTOMATIC1111/stable-diffusion-webui https://github.com/AUTOMATIC1111/stable-diffusion-webui Stable Diffusion https://github.com/Stability-AI/StableDiffusion https://github.com/Stability-AI/StableDiffusion ?
- Aeolun 3y agoThe dog picture is really nice, but then it’s hung on the same wall as 20 other crowded pieces of (in my opinion) dubious quality. This would have been much better standalone.
- renewiltord 3y agoGreat work writing up the process. Much appreciated!
- tonmoy 3y agoIf I wanted to do this, what kind of specs would I need to have on my Desktop Computer?
- cto_official 3y agoIt's impressive, The end result was beautiful. I always use to wonder how to generate some meaningful art Any references where the same has been tried on humans ?
- cto_official 3y agoAlso how much $s were spent on the project?
- frindo 3y agoI did the exact same thing when I saw DreamBooth for the first time! I showed it to a bunch of friends and they convinced me to turn it into an iOS app. https://apps.apple.com/app/ai-avatar-for-dogs-floof-ai/id1659283776 https://apps.apple.com/app/ai-avatar-for-dogs-floof-ai/id165... People have been sending me the cute pics the AI generates of their pups. I think this is arguably the best thing so far in this latest wave of AI releases!
- phonescreen_man 3y agoIs the link an AI generated tutorial? Write me a blog post tutorial pretend you are …
- penthi 3y agoNicely done. I built a t-shirt/mug/frame printing app. I am using stablediffusion (intructpix2pix for selfies) with prompts pulled in from Lexica. The larger images are created with swinir and physical printing is from the good folks at printful.Com. Big props to folks at replicate.Com for making solid infrastructure for ml. https://www.ai-ink.me/ https://www.ai-ink.me/
- deleted 3y ago[deleted]
- mlboss 3y agoAwesome work. I build an app to train dreambooth model and generate images of hich makes this process very easy. The app also has a rest endpoint for anybody to create app using it. Lot of clients create niche websites catering to different use cases. There is a kind of gold rush going on in this area. https://aipaintr.com https://aipaintr.com
- pxoe 3y agothe lengths techbros will go for in order to avoid paying an artist for artwork as well as, doing all that nn/ml stuff, instead of just, trying to learn a bit of how to make an artwork themselves, how to draw something, even by tracing over a photo, like doing a 'how to do a vector colorful painting dog' search and going off on that. like, this end result doesn't even look far off from what a 'colorful vector dog portrait' tutorial would yield. it just involves tons and tons of questionably sourced artwork, and violated copyrights. (i know techbros are very confused about copyrights, but stuff like licenses and copyrights actually do have their meanings, limitations, and liabilities) specifically picking stablediffusion, probably the most blatantly stolen artwork-based model (given how open and clear it is with what data has been used for it, and how you can't just squirm 'i didn't know what were the terms of use of their data' with other, more closed-off services), that's just another great touch as well.
- cztomsik 3y agofeeling better?
- hatefulmoron 3y agowhy would he pay for an artist when he's happy with what he has? Why do drawcels feel so entitled to be looped in financially for no reason? I find the complaint about copyright so strange in this case. Copyright has a purpose, stopping some random person from creating an imagine only they will see and use is not that purpose. In this case it's just spiteful. Ultimately, if you think he's infringing your copyright you should sue him, but I don't think you'd win.
- pxoe 3y agobesides whether artists should get paid or not, or whether they should be reimbursed for use of their art or not, using something without permission, or rights, or license, without something that'd actually (legally) enable to do so without violating copyright, is just bad in itself. there's a great alternative to "not paying/refusing to pay" (but using and stealing stuff anyway) - just, not using other people's stuff. not using stuff that's built on copyright/license violations. not using artwork you don't own, that you don't have rights, licenses, or permissions to use. (yes, simply 'taking something and making a model from it', would be a violation.) one could just not do a shitty thing, and they wouldn't have to jump hoops to find any justification for the shitty thing they did. they could do a step-by-step art tutorial, and wouldn't have to pay anybody, nor use tools that rely on stolen artwork. but nope. highly ironic how they made this thing, and promptly showed it off to thousands of people on the internet, immediately invalidating your example they also promote (just by choosing and mentioning all of these things) those services, like Replicate, that monetize the use of stolen artwork (by selling compute, directly coupled with nn models), and ultimately profit from it (solely, without "giving back to artists whose art they perused" or anything). they could make art in a way that wouldn't participate in tech art theft racket, but they didn't. and they didn't just participate in it, but promote it and perpetuate it.
- ronnykylin 3y agoThe style looks Andy Warhol to me
- mlsu 3y agoFantastic! definitely bookmarking. I spent a big part of the last few days attempting this, my model didn't come out nearly so well. I decided that it was because I don't have enough training images, and so have been taking 3x as many pictures of my dog to compensate.
- dezmou 3y agoI used Stable Diffusion Dreambooth to generate my github profil picture, what do you think ? https://github.com/dezmou https://github.com/dezmou
- Borrible 3y agoRecently saw a nice replica of Duchamp's work 'Bottle Rack' from 1959. Readymade, but maybe a bit expensive. For the price they asked, I could do it myself in a blacksmithing class and have more fun.
- OOPMan 3y ago"art portrait" seems grammatically wrong...?
- blitzar 3y agoThis is the barely a full step from; "I used Stable Diffusion and Dreambooth to create nudes of a person I know". Yes, They're Real and They're Spectacular.
- true_religion 3y agoAll the current stable diffusion commercial applications are in this format: take picture of subject & create a bunch of new portraits of subject. For example: - There's one that creates gaming avatars based on your picture - There's one that makes professional headshots for your CV Even Microsoft's own product (Microsoft Designer) to use AI to create posters & flyers is most useful when you start with an image of your own creation then use the AI to change the style of the image, or integrate it into a template that it dreams up.
- nbzso 3y agoValuable processing info in the comments. But why so much effort to produce something without the option to have ownership (copyright) over my product? If I draw a strange line with any digital painting tool and put a circle and a square around, I sign this art and this is my Art. If a spend a day with prompting, upscaling, fixing with Control net in the end of the day I will have a funny picture which is not mine. https://fortune.com/2023/02/23/no-copyright-images-made-ai-artificial-intelligence/ https://fortune.com/2023/02/23/no-copyright-images-made-ai-a...
- efreak 3y agoSome of us don't care about ownership. You could ask the same question of anyone contributing to open source projects
- nbzso 3y agoThis is a different use case. Why? Because you make a conscious decision to donate your work. The models which are used for image generating (Midjourney, Stable Diffusion) are full of scraped data without consent from the authors. From this point of view, Adobe Firefly is obviously ahead: "The current Firefly generative AI model is trained on a dataset of Adobe Stock, along with openly licensed work and public domain content where copyright has expired. As Firefly evolves, Adobe is exploring ways for creators to be able to train the machine learning model with their own assets, so they can generate content that matches their unique style, branding, and design language without the influence of other creators’ content. Adobe will continue to listen to and work with the creative community to address future developments to the Firefly training models." So the only way forward to have an ownership of your product is to train your own models over your own data.