12 ms·
Running Stable Diffusion on Your GPU with Less Than 10Gb of VRAM
- thepra 4y agoThe issues with f*ng console commands is that they fail, too often. After installing CUDA 11.7 and reinstalling torch I'm still facing: > AssertionError: Torch not compiled with CUDA enabled
- monkmartinez 4y agoI totally understand the frustration. Hop on the Conda train and don't look back. There is no performance penalty from using Conda for the boring stuff. The only thing it will cost you is more disk space. Otherwise, its an absolute joy to use. You know where everything is if you want to inspect packages, bin files, wheels, etc. It seems like chasing your tail when you install these things from apt, git, curl, pip and brew/choco. To me, I want to see where everything has come from and where it is going on my system. Conda gives me that in spades.
- BTCarel 4y agoCan't believe how awesome these generated images are. Thank you for the guide!
- T0Bi 4y agoIf you want to have a really good experience using stable diffusion, use this guide: https://rentry.org/GUItard https://rentry.org/GUItard - includes a nice GUI - txt2img and img2img - upscaling, face correction - many more
- constantlm 4y agoThis is indeed a very thorough, albeit not very nicely named, guide.
- sdflhasjd 4y agoIt originates from 4chan /vg/ & /g/ boards
- hackernewds 4y ago> --ULTIMATE GUI RETARD GUIDE--
- fortyseven 4y agoDo we REALLY this kind of garbage associated with SD? Bad enough I see trashy right-wing extremist shit over on the Stable Diffusion discord server zip past now and then. I'll pass on that guide. Hopefully they grow the fuck up at some point.
- throwaway675309 4y agoUnfortunately as long as there are people who are easily triggered by this sort of thing (seems like they got a bit of a rise out of you) they'll continue in this fashion.
- pavlov 4y agoLet's just pretend it's named after a background process that keeps track of your guitar.
- Datagenerator 4y agoCan this be used with the optimizedSD by basujindal?
- mutant 4y agoEdgy title
- mugivarra69 4y agoanyone tried to quantize or use bfloat?
- qayxc 4y agoblfoat would indeed be nice. It's supported on a wide range of hardware (basically all mid-range to high-end Intel CPUs since 2013, AMD MI5 and up compute cards, ARM NEON and NVIDIA cards since Pascal [10-series, 2016!]). It could speed up calculations and significantly reduce memory requirements. I'd expect slightly worse results, though. edit: also https://github.com/basujindal/stable-diffusion/pull/103 https://github.com/basujindal/stable-diffusion/pull/103
- mugivarra69 4y agoneat. thanks!
- deleted 4y ago[deleted]
- cube2222 4y agoFor those without a GPU / not a powerful enough one / wanting to use SD on the go, you can start the hlky stable diffusion webui (yes, web ui) in Google Colab with this notebook[0]. It's simple and it works, using colab for processing but actually giving you a URL (ngrok-style) to open the pretty web ui in your browser. I've been using that on-the-go when not at my PC and it's been working very well for me (after trying numerous other colab-dedicated repos, trying to fix them, and failing). Additionally, you can have all your generated images sync to Google Drive automatically. [0]: https://github.com/altryne/sd-webui-colab https://github.com/altryne/sd-webui-colab
- Llamamoe 4y agoWho's paying for all the Google Collab notebooks I've been seeing around? Can I really just start and keep using it for free?
- Karuma 4y agoGoogle is paying, and yes, you can, but they will disconnect you after a while. And if you abuse it too much, you won't be able to use it until the following day... You can also buy Colab Pro and Colab Pro+, which have fewer limitations and faster GPUs.
- capableweb 4y agoHow fast is the Colab stuff? Is Colab Pro/Pro+ a lot faster too? I run it locally and can generate images with 50 steps in about 6 seconds per image, would it be faster for me to use Colab Free/Pro/Pro+?
- cube2222 4y agoIn my usage Colab and Colab Pro were similar, with plain Colab occasionally OOMing during model loading. That said I've actually been seeing times slower than yours on Colab and I think they're slower than on my RTX 3080. ~15 secs per image. I'm not sure why, though.
- blfr 4y agoIs there a similar guide for Linux/Ubuntu with some sort of light sandboxing, at least python virtual virtual environment?
- politelemon 4y agoYes I followed one recently, though it uses conda. The SD script runs in a conda environment, so when you uninstall conda your system is preserved and hasn't been stomped on. https://code.mendhak.com/run-stable-diffusion-on-ubuntu/ https://code.mendhak.com/run-stable-diffusion-on-ubuntu/
- blfr 4y agoLike you read my mind. Thank you!
- forgingahead 4y agohttps://github.com/basujindal/stable-diffusion https://github.com/basujindal/stable-diffusion I use this on my Ubuntu 18 machine, works nicely on a GPU with 8GB VRAM. As usual, some python dependency nonsense to sort out even with Anaconda, but pretty quick and easy to get up and running.
- Datagenerator 4y agoThis one is from scratch on Debian: https://notes.datagenerator.eu/#Stable%20Diffusion%20installation%20on%20Debian%2011%20(bullseye) https://notes.datagenerator.eu/#Stable%20Diffusion%20install...
- DarthNebo 4y agoAlways wondered why we can't virtualize VRAM like how we did for VMs.
- kernelsanderz 4y agoYou kind of can - projects like deepspeed (https://www.deepspeed.ai/ https://www.deepspeed.ai/) enable running a model that is larger than in VRAM through various tricks like moving weights from regular system RAM into VRAM between layers. Can come with a performance hit though depending on the model, of course.
- WithinReason 4y agoGood question. Bandwidth of dual channel DDR4-3600: 48 GB/s Bandwidth of PCIe 4 x16: 26 GB/s Bandiwdth of 3090 GDDR6X memory: 935.8 GB/s Since neural network evaluation is usually bandwidth limited, it's possible that pushing the data through PCI-E from CPU to GPU is actually slower than doing the evaluation on CPU only for typical neural networks. https://www.microway.com/knowledge-center-articles/performance-characteristics-of-common-transports-buses/ https://www.microway.com/knowledge-center-articles/performan... https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_processing_units#GeForce_30_series https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces...
- redox99 4y agoAnd that's without even taking into account latency of accessing main memory through PCIe, which would make matters even worse.
- sp332 4y agoOk, but at least it would run.
- zamadatix 4y agoWhat's the point of running it on the GPU if to do so you need to make it slower tham running in the CPU? Just run it on the CPU at that point.
- reckless 4y agoWould be great to be able to utilise outpainting to generate larger images in smaller tiles at full precision.
- mabbo 4y agoI believe I saw a repo that was doing exactly that. They also included a step at the end to reintegrate the results better. I was also able to use the basic scripts to generate a few samples, pick one I liked, then used inpaint to expand the photo, masking out the original input so it wouldn't be altered.
- barrkel 4y agoI started out using my old GTX 1080 on Thursday, could generate 512x512 just fine. That's in 8G of VRAM. It worked well on the hlky branch using webui (built using gradio). Seeing that training etc. is much more memory intensive, and wanting to get faster results, I bought an RTX 3090, which has 24G of VRAM. However it maxes out at about 1024x512, only twice as many pixels. Observing the card with GPUZ, it never actually allocates more than 13.9G. Using the lstein branch, I can't get above 896x512. Similarly, GPUZ shows allocated VRAM never reaches 14G. The interface isn't as good as the webui on hlky either - never mind the web interface, a bigger problem is it doesn't save all the parameters alongside generated images. This is all running using Miniconda on Windows. On Linux it may be a different story, but my gaming PC is not dual-boot (yet).
- WhereWhyWhat 4y agoHow many it/s do you get with 3090 compared to the 1080? My 1080 Ti gets around 2.5it/s with the k_lms sampler.
- capableweb 4y agoOn a 2080 Ti I get around 8 it/s with k_lms
- barrkel 4y agoWhere the 1080 would do about 2 it/sec, the 3090 does about 10 it/sec. When I do batches, it slows down but not linearly; if 1x does 10 it/sec, 2x does about 6 it/sec. Batching is the other upside of more VRAM.
- pixl97 4y agoMaybe look in to booting off a USB stick as a means to test this. I wouldn't be surprised if there were some kind of driver reservation I Windows causing this issue.
- MattRix 4y agoHave you tried something like 1024x768? Going to full 1024x1024 would double your VRAM usage so I can see why that wouldn’t work. For my uses, the real benefit of having more VRAM is that you can generate more images simultaneously. My 3080 can generate only one 512x512 in 7 seconds but three 384x384 in that same timeframe. It’s allowed me to generate grids of hundreds of images in just a few minutes.
- lagrange77 4y agohttps://constant.meiring.nz/assets/posts/2022-08-04-playing-with-stable-diffusion/A%20cat%20smoking%20a%20cigarette.%20Moody.%20Dark.-6422.jpeg https://constant.meiring.nz/assets/posts/2022-08-04-playing-... How it holds the cigarette with its little paw. Ehem, i mean, it's technically interesting, how the model correctly extrapolated, how this would look like..
- alkonaut 4y agoWhat’s the easiest way of using SD on a Windows box? Can I run it off a Linux live USB or can it run directly under Windows? Edit: never mind this is the missing guide I had been looking for
- constantlm 4y agoThe guide posted is for Windows 11.
- andybak 4y agoOn Windows the app Visions of Chaos (mostly)-automates the installs for dozens of ML models including SD: https://softology.pro/tutorials/tensorflow/tensorflow.htm https://softology.pro/tutorials/tensorflow/tensorflow.htm and provides a fairly respectable UI. It's also updated almost daily and tracks the latest features where possible.
- redacted 4y agoThe Linux/not-Windows instructions on https://github.com/hlky/stable-diffusion/wiki/Docker-Guide https://github.com/hlky/stable-diffusion/wiki/Docker-Guide worked well for me using WSL2 with nvidia-docker
- hwers 4y agoIf you have even just 4gb stable diffusion will run fine if u go for 448x448 instead (basically the same quality).
- SuperCuber 4y agoI feel like I'm going insane. Everyone says 512x512 should work with 8gb but when I do it I get: CUDA out of memory. Tried to allocate 3.00 GiB (GPU 0; 8.00 GiB total capacity; 5.62 GiB already allocated; 0 bytes free; 5.74 GiB reserved in total by PyTorch) any ideas? I have a 3060ti with 8gb vram... with 448x448 I get: CUDA out of memory. Tried to allocate 902.00 MiB (GPU 0; 8.00 GiB total capacity; 6.73 GiB already allocated; 0 bytes free; 6.86 GiB reserved in total by PyTorch)
- schleck8 4y agoUse halfprecision float and/or the optimized forks https://github.com/basujindal/stable-diffusion https://github.com/basujindal/stable-diffusion https://github.com/neonsecret/stable-diffusion https://github.com/neonsecret/stable-diffusion Or the hlky webui, that is optimized too. http://rentry.co/kretard http://rentry.co/kretard
- baobabKoodaa 4y agoI've been trying to get the basujindal fork to work, but it seems to be putting all work on the CPU. I've been running the example txt2img prompt for 30 minutes now and it's still not finished. It has reserved 4Gb memory from the GPU, but the GPU doesn't appear to be doing any work, only CPU is doing work.
- deleted 4y ago[deleted]
- prettydeep 4y agoUse the original SD repo. But modify the txt2img.py according to: https://github.com/CompVis/stable-diffusion/issues/86#issuecomment-1230309710 https://github.com/CompVis/stable-diffusion/issues/86#issuec...
- baobabKoodaa 4y agoIt's unfortunate that this article doesn't specify the amount of VRAM needed, other than specifying it's "less than 10Gb". I have 6,1Gb of VRAM and I tried to follow the article until eventually encountering an "unable to allocate memory" error. (I'm now trying to run basujindal's repo as an alternative.)
- capableweb 4y agoReduce the resolution and run with half-precision instead of full-precision and you should be able to avoid OOM errors. Author seems to have had 8GB VRAM available, so I'm guessing that's the "minimum required" for their solution.
- baobabKoodaa 4y agoIt's not possible to halve the precision further. The precision was already dropped from float32 to float16 in the OP. I now used parameters to drop the resolution to 256x256, and now it's running, but it's somehow broken. Every output image it produces is literally a green square.
- LanternLight83 4y agoThe green square issue has been well known, particularly on AMD cards, and I believe the solution is... full precision :c But idk, I haven't had that issue. My issue's that I can run it in <4GB VRAM, but can only do a couple dozen images before some memory leak or smth drives it out of memory (effects my 2070S too, but only after many more images). Restarting it isn't too bad, but it's enough to have me looking to using either if two AMD APU's that I have on hand.
- XorNot 4y agoYou need to be in full precision mode in that case. Running on my AMD card this was necessary.
- baobabKoodaa 4y ago
- verytrivial 4y agoFrom the diff, perhaps stale but: > Carbon Emitted (Power consumption x Time x Carbon produced based on location of power grid): 11250 kg CO2 eq. That's ... Sobering.
- IshKebab 4y agoMy work uses a monorepo without precise dependency tracking (Bazel or similar) so every single diff builds everything and runs a ton of tests. About 6 kWh of electricity per diff. Even for typos. Nobody seems especially bothered.
- XorNot 4y agoI have this running on my fairly mundane Radeon 5600XT at about 1 minute per image generated (under rootless podman, which is the real cool news to me) which isn't bad all things considered. Definitely get some interesting sounds from coil whine when it's going.
- nilolo 4y agoCould you share which repo you got running? I followed this guide (https://gitgudblog.vercel.app/posts/stable-diffusion-amd-win10 https://gitgudblog.vercel.app/posts/stable-diffusion-amd-win...) that uses Onnx to get it running on my 5700XT but I'm not happy with the performance at about 2m30s per image.
- mtoddsmith 4y agoIntegrate this into a game for infinite playability. Does the image generator return some kind of seed that allows you to reproduce the result?
- the-golden-one 4y agoThe seed is passed on the command line.
- hedora 4y agoI had good luck with these directions, which let you run inside a docker container: https://github.com/AshleyYakeley/stable-diffusion-rocm https://github.com/AshleyYakeley/stable-diffusion-rocm I had to make the one line change suggested in issue #3 to get it to run under 8GB. radeontop suggests 4GB might work. I also had to add this environment variable to make it work on my unsupported radeon 6600xt: HSA_OVERRIDE_GFX_VERSION=10.3.0 It takes under two minutes per batch of 5 images with the --turbo option. (Base OS is manjaro; using the distro's version of docker; not the flatpack docker package.) If you don't have a GPU, paperspace will rent you an appropriate VM.
- SubiculumCode 4y agoAll I keep thinking is, how can I make money off of this. Aww the power of open-source. Right now, my thinking is that its just going to cut costs (sorry artists) for in existing workflows, maybe change some endeavors from red to black profit margins. Probably more likely will be using this SD as a basis for more specialized content training.
- spapas82 4y agoI'd like to confirm that this works in my GTX 2060 with 6 GB VRAM on windows. I didn't do any modifications on the provided source code; faces are a little problematic. I don't use anaconda so I created a new venv with python 3.10, installed the requirements as proposed, registered with hugging face and create the api key and run the provided source code. Any way to improve the quality of the faces? Also how could I tune the parameters a bit ? (I'm not familiar with this AI stuff at all, I'm just a humble python programmer)
- r2_pilot 4y agoIt's my understanding that v1.5 will be coming out in a few weeks; I recall that hands and faces will be better-trained in the new model. I'm about to try what you did (install requirements manually) to get it to run in PyCharm on Windows. Neither miniconda nor anaconda really worked for me after spending time trying to get it to pick up the dependencies.
- password4321 4y agoI didn't realize 512x512 on 4GB VRAM (Win10 over RDP) was anything unusual, just followed https://github.com/awesome-stable-diffusion/awesome-stable-diffusion https://github.com/awesome-stable-diffusion/awesome-stable-d... to "Optimized Stable Diffusion" https://github.com/basujindal/stable-diffusion https://github.com/basujindal/stable-diffusion (linked many times in this discussion).
- cdelsolar 4y agoanyone know how to get conda running on arch linux? `conda init bash` gives me some Python errors.
- moron4hire 4y agoYeah, I don't think that's an Arch Linux problem. I had similar problems on Windows, and one version of the project even was supposedly setup to run in Docker. What is the point of setting up Docker if the whole setup and build process is not turnkey? Seems like all of these projects are broken until you speak shibboleth by guessing at random python incantations. By this point, it's starting to feel intentional, like a way to mark you as part of an in-crowd, not a "L-User". Unfortunately, I don't remember what I did. I did eventually get SD to work (though not in Docker, just as a normal python project). If I had been sober at the time, I probably would have given up. I know you need no greater than Python 3.9.
- cdelsolar 4y agoI'm on Python 3.10, so I probably have to wait till they fix something.
- moron4hire 4y agoYeah, I don't really get why Python developers tolerate this state of being where minor point releases are so commonly incompatible with each other, or having no standard way to manage building against different versions. I also don't get why ML developers continue with Python, considering these issues. But, that's the point of the Miniconda dependency. By using Conda, the project can be setup locally with the Python version it expects without clobbering your local, system-level install. It was still a bit of a pain to learn how to use Conda, but it worked out a little better than figuring it out on my own.
- neurostimulant 4y agoPretty easy actually. Just install miniconda from here [1]. It'll add some codes into your bashrc / zshrc so you'll need to reopen your terminal after installation. [1] https://docs.conda.io/en/latest/miniconda.html#linux-installers https://docs.conda.io/en/latest/miniconda.html#linux-install...
- layer8 4y agoI’ll get downvoted, but it’s a genuine question: Will “a photo of tits and ass” generate photos of birds with donkeys, or will it rickroll you [0]? [0] https://twitter.com/qDot/status/1565076751465648128 https://twitter.com/qDot/status/1565076751465648128
- lbotos 4y agoDepends on if you are running stable diffusion with the safety filter on or not. By default it's on, some forks have it turned off.
- lbotos 4y agoAs I understand it, SD was trained on this dataset: https://rom1504.github.io/clip-retrieval/?back=https%3A%2F%2Fknn5.laion.ai&index=laion5B&useMclip=false https://rom1504.github.io/clip-retrieval/?back=https%3A%2F%2... So go here, turn off the safety filter and you can search to see what SD was trained on. I suspect that if you actually want the bush tit bird and donkeys, you'll want to use that instead.
- coolspot 4y agoJust tested it locally[0] with two prompts (all default params): "A photo of tits and ass" and "A photo of tits (birds) and ass (donkey)" Result: https://imgur.com/a/c1GM28U https://imgur.com/a/c1GM28U (NSFW) Censored version would just replace anything NSFW with a picture of Rick Astley (for real [1]). [0] - https://github.com/hlky/stable-diffusion https://github.com/hlky/stable-diffusion [1] - https://github.com/CompVis/stable-diffusion/issues/120 https://github.com/CompVis/stable-diffusion/issues/120
- layer8 4y agoThanks for actually trying that out. Surprisingly few tits with those asses. The (animal) ass-tit chimeras are amusing.
- hombre_fatal 4y agoI've been running Stable Diffusion on my M1 Macbook since the thread a few days ago about doing just that. I am comically bad at getting it to generate what I want. e.g. "A furry watermelon" or "A dog flexing its biceps" just generates normal watermelons and normal dogs most of the time. Any tips?
- nickthegreek 4y agoUse www.lexica.art for prompts that work, modify the keywords for your needs.
- davidy123 4y agoNot an Apple guy, but I think an Apple M chip will run at ⅓ the speed of a top end RTX GPU, however it uses system memory, so it can easily be 32GB or 64Gb. That's pretty compelling, and if this is really a new class of application, NVidia is going to have to think about more memory for mainstream-ish products.
- skybrian 4y agoThis is a specialty application. I don't think it's going to be big enough to drive consumer technology like gaming? Particularly since cloud services are likely to be competitive and work for anyone.
- coolspot 4y agoIt is 50x times slower on M1 than on RTX 3090. M1 takes ~4.2s per iteration, 3.5 minutes per image [0]. RTX 3090 takes ~4.7s per image (all 50 iterations) [1]. [0] - https://wandb.ai/morgan/stable-diffusion/reports/Running-Stable-Diffusion-on-an-Apple-M1-Mac-with-Hugging-Face-Diffusers--VmlldzoyNTU2ODc2 https://wandb.ai/morgan/stable-diffusion/reports/Running-Sta... [1] - trust me bro
- shrimpx 4y agoBtw that's the kind of perf I see on my M1, but I keep seeing "0.00G VRAM used" for each generation. I wonder what that's about. In Activity Monitor I do see the GPU being used.
- coolspot 4y agoSD measures VRAM usage by calling a specific pyTorch method which usually wraps CUDA call. I guess whomever ported that to M1 just haven’t implemented that method.
- davidy123 4y agoOK, I must have misread some comments. Thanks for the update.
- figomore 4y agoOther option is to use the Openvino one (https://github.com/bes-dev/stable_diffusion.openvino https://github.com/bes-dev/stable_diffusion.openvino). It uses CPU and runs very fast. It takes ~90s to generate an image on my Ryzen 3800X.
- Joyfield 4y agoGb != GB.
- constantlm 4y agoThank you - fixed it on the post title. Not sure it'll update on here though.