12 ms·
Stable Diffusion PR optimizes VRAM, generate 576x1280 images with 6 GB VRAM
- MayeulC 4y agoI wonder if GTT (GART size) wouldn't help with limitted VRAM amounts, at the cost of some performance?
- avocado2 4y agocan run with gui: https://github.com/basujindal/stable-diffusion/blob/main/optimizedSD/txt2img_gradio.py https://github.com/basujindal/stable-diffusion/blob/main/opt...
- jacobn 4y agoInteresting that “del” helps here - could eg torch.jit.script/trace be more of a “sledgehammer” solution?
- ceeplusplus 4y agoUnfortunately the author of this PR decided to include garbage like this [1] in it, so this PR is pretty useless. [1] https://github.com/basujindal/stable-diffusion/pull/103/files#diff-c693279643b8cd5d248172d9c22cb7cf4ed163a3c98c8a3f69c2717edd3eacb7 https://github.com/basujindal/stable-diffusion/pull/103/file...
- deleted 4y ago[deleted]
- kevingadd 4y agoYou're still able to run it locally even if the license change prevents it from being merged, and it seems to be a proof of concept that may inspire other people to optimize it in different ways
- Matheus28 4y agoSeems like author did his work on a different fork and pushed to this one, which included all changes from the other fork...
- jinseokim 4y agoI just reminded "clean room design" technique: https://en.m.wikipedia.org/wiki/Clean_room_design https://en.m.wikipedia.org/wiki/Clean_room_design
- javchz 4y agoYeah, reminds me of the old freedoom WADs situation
- cercatrova 4y agoHow is that garbage? That license is part of the original SD repository, the creator Emad even talks about it in his initial post, about the OpenRAIL M License [0]: > i) The model is being released under a Creative ML OpenRAIL-M license [https://huggingface.co/spaces/CompVis/stable-diffusion-license https://huggingface.co/spaces/CompVis/stable-diffusion-licen...]. This is a permissive license that allows for commercial and non-commercial usage. This license is focused on ethical and legal use of the model as your responsibility and must accompany any distribution of the model. It must also be made available to end users of the model in any service on it. [0] https://stability.ai/blog/stable-diffusion-public-release https://stability.ai/blog/stable-diffusion-public-release
- deleted 4y ago[deleted]
- zarzavat 4y agoIt’s worth to point out that there’s no consensus yet on whether ML models (as in the weights) are actually copyrightable. If you’re using them, you should probably assume that they are. If you’re distributing them, you should probably assume that they aren’t.
- whywhywhywhy 4y agothey mean the entire thing has been run through some sort of "beautifier" so there are thousands of irrelevant whitespace changes across things unrelated to the commit, which itself is around 6 lines.
- sercand 4y agoThis is the worst type of PR. Just submit a different PR for the code format changes.
- dr_dshiv 4y agoHow much VRAM was previously required?
- temp_account_32 4y agoI am able to generate a maximum resolution of 512x768 on my 11GB 1080Ti. This seems to use almost 100% of the available RAM.
- frognumber 4y agoThat's weird. I cap out at 512x512 on my 16GB Ampere card. Even stepping down precision doesn't help. I wonder what's different. I use it directly from Python.
- prox 4y agoPerhaps a different kind of chipset that isn’t optimized yet? Just a guess based on other graphics processes.
- MattRix 4y agoYou might have the n_samples (aka batch size) set to a number greater than 1? That basically multiplies the amount of VRAM you’re using. I can generate a 512x512 on my 10gb 3080 no problem (or three 384x384 at a time)
- temp_account_32 4y agoTry the hlky fork: https://github.com/hlky/stable-diffusion https://github.com/hlky/stable-diffusion It measures memory usage as well.
- gpt5 4y agoI can generate a 768x896 pixels image on an RTX 3090, using 23.4/24GB
- moneycantbuy 4y ago
- oefrha 4y agoAlmost impossible to pinpoint what changed thanks to thousands of lines of completely irrelevant changes and shitty commit messages. It seems the only changeset that might be relevant out of +2,273 -1,531 is the +11 -7 from https://github.com/basujindal/stable-diffusion/pull/103/commits/47f878421c5bf97d0fff44edaa703d152cafb483 https://github.com/basujindal/stable-diffusion/pull/103/comm...? Does it even work?
- piyh 4y agoAs a learning opportunity for people like me, what does a good PR look like for a large change?
- adastra22 4y agoYou don’t make large change PRs.
- MontagFTB 4y agoBranching in git is very cheap, so the typical path a large change takes to the main branch is by a series of smaller, incremental PRs. This keeps your work close to the main branch as much as possible.
- oefrha 4y agoI would say large contributions from non-members generally work rather terribly in open source. If you absolutely have to, - Communicate ahead of time; don't surprise maintainers with sweeping architectural changes or huge features no one wants or would like to review; - Try to break up changes into logical units that can be understood and reviewed independently; - Write useful and detailed commit messages (some bad examples: "Update attention.py", "various clean-ups, code now beautified"); - Don't sneak in anything unrelated to the PR; don't sneak anything unrelated into a commit; - Absolutely don't use a code formatter to format the entire code base if the repo wasn't already using one. You can suggest that separately. And changes like that are best done by a trusted member.
- zo1 4y ago
- qdot76367 4y agoFor everyone about to comment on the garbage in the commit: It looks like the committer made their changes in the top commit, then merged the updated CompViz StableDiffusion change set on top of it for some reason. That's where the license change, rick astley image, etc come from. And yes, StableDiffusion from the original repo will rick roll you if you try to generate something that triggers its NSFW filter. Here's the code that does it: https://github.com/CompVis/stable-diffusion/blob/main/scripts/txt2img.py#L79 https://github.com/CompVis/stable-diffusion/blob/main/script... And here's what it looks like: https://twitter.com/qDot/status/1565076751465648128 https://twitter.com/qDot/status/1565076751465648128
- NHQ 4y agoBut what is the correct git command to ignore all that?
- cercatrova 4y agoThere's no git command to ignore it, the main repo should merge their changes and then the other person can make a PR on the updated changes.
- NHQ 4y agoWhere are PRs in Git?
- oefrha 4y agoDo a git-cherry-pick.
- p-e-w 4y ago> And yes, StableDiffusion from the original repo will rick roll you if you try to generate something that triggers its NSFW filter. It goes without saying that the authors of a piece of software have the right to make the software do whatever they want, but that shouldn't stop us from recognizing that AI engineers are starting to act like megalomaniac overseers who consider it part of their mission to steer humanity onto the "right" path. Who exactly do these people think they are? Imagine this behavior from a web browser. "The URL of the file you were trying to download triggered my NSFW classifier, so I'm going to replace the file with this funny image." This isn't funny, it's creepy.
- deleted 4y ago[deleted]
- nl 4y agoI've been using the HuggingFace diffuses repo[1] with 6GB of VRAM fine. It's well engineered, maintainable and with decent installation process. The branch in this PR[2] adds M1 Mac support with a one line patch and it runs faster than the CompVis version (1.5 iterations/sec vs 1.4 for CompVis on a 32 Gb M1 Max, I highly recommend people switching to that version for the improved flexibility. [1] https://github.com/huggingface/diffusers https://github.com/huggingface/diffusers [2] https://github.com/huggingface/diffusers/pull/278 https://github.com/huggingface/diffusers/pull/278
- JimDabell 4y agoI wouldn’t recommend using that as-is. MPS doesn’t give deterministic random number generation, which means that seeds become meaningless and you won’t ever be able to reproduce something. You can work around it by generating random numbers on the CPU and then moving them to MPS, but that probably requires a fix in PyTorch. The MPS support issue for diffusers is here: https://github.com/huggingface/diffusers/issues/292 https://github.com/huggingface/diffusers/issues/292 …and it links to the relevant PyTorch issue here: https://github.com/pytorch/pytorch/issues/84288 https://github.com/pytorch/pytorch/issues/84288
- nl 4y agoThat's really annoying, even though I hadn't noticed it until now. I can confirm this is an issue. But the situation seems to be the same on the CompVis derived repos, right? So this is no worse off, but with better engineered and faster code.
- JimDabell 4y agoYeah, but with the CompVis derived repos, it’s pretty easy to go in and change all the calls to PyTorch random number generators. Having said that, the last comment [0] on the PyTorch issue gave me the idea of monkey patching the random functions. The supplied code assumes you’re always passing in a generator, which is not true in this case, but if you monkey patch the three rand/randn/randn_like functions to do nothing but swap out the device parameter for 'cpu' and then call to('ops') on the return value, it’s enough to get stable seed functionality for the CompVis derived repos without modifying their code, so I’m guessing it will probably work for diffusers as well. Also, it’s probably a bug in the CompVis code, but even after you fix the random number generator, the very first run in a session uses an incorrect seed. The workaround is to generate an image once to throw away whenever you start a new session. [0] https://github.com/pytorch/pytorch/issues/84288#issuecomment-1236247109 https://github.com/pytorch/pytorch/issues/84288#issuecomment...
- andrethegiant 4y agoWild how fast this is moving
- bongobingo1 4y agoJust wait until the AI is writing the code too. You wont be able to `git pull` fast enough.
- jpeter 4y agoThe power of open source
- switchers 4y agoSlightly off tangent, has anyone got it running on an AMD card? Sad 6600xt Windows user reporting in :(
- universa1 4y agoNot sure about the 6600, but there is a guide for Linux at least: https://m.youtube.com/watch?v=d_CgaHyA_n4&feature=emb_logo https://m.youtube.com/watch?v=d_CgaHyA_n4&feature=emb_logo And this is somehow relevant (possibly), as I kept the link open. https://github.com/RadeonOpenCompute/ROCm-docker/issues/38 https://github.com/RadeonOpenCompute/ROCm-docker/issues/38
- mpaepper 4y agoHere is a guide for AMD, but I don't have such a card, so haven't tried. https://rentry.org/sdamd https://rentry.org/sdamd
- saurik 4y agoI was under the impression that generating images where both dimensions were larger than 512 didn't merely require a lot of resources but didn't work well as the model was trained exclusively on 512 by 512 images and while it sort of worked ok to stretch one dimension a bit you got weird repetitions by making the overall canvas too large (as this isn't going to generate higher dpi images, merely ones with greater square inches).
- prettydeep 4y agoThis is accurate in my experience. Changing the res. to anything but 512x512 produces inferior results.
- TaylorAlexander 4y agoIn case anyone is confused by the clashing repos, here is how I was able to easily run this updated code. Clone the original SD repo, which is what this code was built off of, and follow all the installation instructions: https://github.com/CompVis/stable-diffusion https://github.com/CompVis/stable-diffusion In that repo, replace the file ldm/modules/attention.py with this file: https://raw.githubusercontent.com/neonsecret/stable-diffusion/main/ldm/modules/attention.py https://raw.githubusercontent.com/neonsecret/stable-diffusio... Now run a new prompt with a larger image. Note that the original model was trained on 512x512 and may lead to repetition especially if you try to increase both dimensions (this is mentioned in the SD readme) so just run with one dimension increased. For example try the following example: python scripts/txt2img.py --prompt "a person gardening, by claude monet" --ddim_steps 50 --seed 12000 --scale 9 --n_iter=1 --n_samples=1 --H=512 --W=1024 --skip_grid I confirmed that if I run that command with the original attention.py, it fails due to lack of memory. With the new attention.py, it succeeds. That said, this still uses 13GB of ram on my system. I suppose you can check out the full repo with the updated code, which seems to have other changes, if you want to give that a try. https://github.com/neonsecret/stable-diffusion/ https://github.com/neonsecret/stable-diffusion/ I have already been using the original SD repo so I found benefit by just changing attention.py
- macrolime 4y agoI've seen people claim that it's better to generate 512x512 images since that's what it's trained on, and then upscale, rather than generating higher resolution images directly. Anyone tried any systematic investigation of this?
- lionkor 4y agoThis is one of the worst PRs i've ever seen. If I was involved, this PR would be closed and conversation locked, simply because its a ton of formatting changes. The only "memory efficiency" changes are like 6 lines of python where the author uses `del`. Thats it. Everything else is formatting bullshit, and some other merge stuff. Yikes. Also, some of those PR comments are written by kids or something. Wtf
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- WithinReason 4y agoThe diff that apparently does the RAM optimisation, quite simple and something to learn form: https://github.com/basujindal/stable-diffusion/commit/47f878421c5bf97d0fff44edaa703d152cafb483 https://github.com/basujindal/stable-diffusion/commit/47f878...
- liuliu 4y agoYeah, stable diffusion's PyTorch code is not optimized for inference memory usage at start. I am looking at the code now and it seems that if it is converted to static graph, there are probably a bit more opportunities (I only looked at CLIP model and UNet model it uses today, not sure about the Autoencoder yet).
- LarsDu88 4y agoStarted using this branch and it works on my old ass 1060 gpu. Maybe I'll do a blog post on the whole process