31 ms·
Run Stable Diffusion on Your M1 Mac’s GPU
- ThrowawayTestr 4y agoThat was fast.
- omginternets 4y agoIs there a way to get it to run on an an Intel-based Mac? I've attempted several times, but quickly ran into dependency issues and other quirks.
- jw1224 4y agoI believe the branch which adds support for Apple Silicon also adds support for running on Intel chips (albeit extremely slowly). I haven't tested it myself, but I've seen several people in the GitHub issues saying this.
- pja 4y agoThe standard release works fine, if you tweak the code to use the CPU pytorch device instead of CUDA. It does take about an hour to generate a set of images with the standard options on my AMD 2600 CPU though!
- vimy 4y agoComment from github: "By the way, i confirmed to work on my Intel 16-in MacBook Pro via mps. GPU (Radeon Pro 5500M 8GB) usage is 70-80% and It takes 3 min where --n_samples 1 --n_iter 1. My repo https://github.com/cruller0704/stable-diffusion-intel-mac https://github.com/cruller0704/stable-diffusion-intel-mac" For comparison, my RTX 2070 takes 10 seconds for one image (512x512)
- amelius 4y agoI'd rather see someone implemented glue that allows you to run arbitrary (deep learning) code on any platform. I mean, are we going to see X on M1 Mac, for any X now in the future? Also, weren't torch and tensorflow supposed to be this glue?
- nathas 4y agoBroadly speaking, it looks like they are. The implementation of Stable Diffusion doesn't appear to be using all of those features correctly (i.e. device selection fails if you don't have CUDA enabled even though MPS (https://pytorch.org/docs/stable/notes/mps.html https://pytorch.org/docs/stable/notes/mps.html) is supported by PyTorch. Similar goes for quirks of Tensorflow that weren't taken advantage of. That's largely the work that is on-going in the OSX and M1 forks.
- davedx 4y agoI got stuck on this roadblock, couldn’t get CUDA to work on my Mac, was very confusing
- desindol 4y agoDidn’t apple stop supporting Nvidia cards like 5 years ago? How could it be confusing that Cuda wouldn’t run?
- root_axis 4y agolol presumably the OP didn't know that... hence the confusion.
- davedx 4y agoAh, I didn't realize. It's not very obvious what GPU you have in your Macbook, I couldn't actually find out where to find that in my System settings. On Windows it's inside the "Display" settings but on MacOS... where is it? :)
- cercatrova 4y agoThat's because CUDA is only for Nvidia GPUs and Apple doesn't support Nvidia GPUs, it has its own now.
- dustingetz 4y ago(base) stable-diffusion git:(main) conda env create -f environment.yaml Collecting package metadata (repodata.json): done Solving environment: failed ResolvePackageNotFound: - cudatoolkit=11.3 oh i was following the github fork readme, there is a special macos blog post
- keepquestioning 4y agoGamechanger!
- keepquestioning 4y agoOne beautiful thing I realized about all this progress in AI. We will still need people to do the hard yards, and get dirt between their fingernails. I am firmly in the camp of those people. Fancy algorithms won't dig holes, or lay out rail tracks of over hundreds of miles.. or build houses all across the world.
- blagie 4y agoAre you following progress in robotics?
- dougmwne 4y agoSchool us! What’s the latest in robotics that is going to knock our socks off?
- scoopertrooper 4y agoThey've got death robots that can fly now. Why is nobody impressed by the future?
- dougmwne 4y agoI believe those are human controlled, no? Robotics gets really interesting when the robots can start driving and building the roads as well.
- neurostimulant 4y agoNot necessarily. Some just require human to turn it on and it'll loiter and attack enemy autonomously (loitering munition [1]). [1] https://en.wikipedia.org/wiki/Loitering_munition https://en.wikipedia.org/wiki/Loitering_munition
- blagie 4y agoI mean, I'm most impressed with the gradual rise of 3d printers capable of printing entire houses. That doesn't help with plumbing, electrical, kitchen cabinets and all the other stuff that goes into a house -- which is a majority of the cost -- but it's a gradual start. I see very little which can't be automated, but I see a lot which would take many years of time and effort to automate and integrate.
- avereveard 4y agoHow fast is it on a m1?
- ebiester 4y agoIt takes 1-2 mins for a 512x512 image. It's been a lot of fun since I did this last night.
- smoldesu 4y agoFor reference, inferencing the model on a 2070 takes 10-12 seconds for the same size at max-precision, and the 3070 can synthesize an image in almost 6 seconds. If you extrapolate the power consumption (3070 @~300w vs M1 Pro GPU@~30-50w) the metrics make a lot of sense.
- nathas 4y agoI haven't ran this fork yet, but about 1.3 sec/iter. Usually ~30-50 iters/sample (image).
- pwinnski 4y agoThe answer depends VERY MUCH on RAM. My M1 with 8GB takes 70-90 minutes per image. My M1 Pro with 16GB takes 3 minutes per image.
- ebiester 4y agoNote: I ran this and haven't yet been able to get img2img working yet. I borked it up trying to get conda working. It's been a lot of fun to play with so far though!
- bravura 4y agoApparently, this release should include a Dockerfile for easier replicability.
- bfirsh 4y agoUnfortunately this can't run in Docker because Docker for Mac can't access the M1 GPU. (Several layers of virtualization and emulation!)
- nathas 4y agoTry the lstein fork: https://github.com/lstein/stable-diffusion/tree/fix-cuda-reset-stats-error https://github.com/lstein/stable-diffusion/tree/fix-cuda-res... You'll still need to play with modifying some of the code to get it to run, but `dream.py` works for me. Funny enough, I got only img2img effectively working with the lstein branch; it broke txt2img for me.
- StapleHorse 4y agoYesterday I thought I broke it too. In my case, the solution was just to make sure that the input image (from the editor or otherwise) was the same size as the output image. Hope it helps.
- caxco93 4y agocould someone who has already done this please share how long it takes for a 50 steps image to be generated?
- nathas 4y ago1.3 sec/iter on my M1 Mac, so ~39 seconds.
- jw1224 4y agoThat was fast. I'm only getting 5.26s/iter on an M1 Pro MBP with 16GB RAM. EDIT: Speed increased to 2.3s/iter after a reboot
- nathas 4y agoDepends what fork you're running... Some seem to be using CPU-based generation, others use the MPS device backend correctly which is MUCH faster. I have another comment floating around about lstein's fork, but it takes some massaging to get it to run happily. https://github.com/lstein/stable-diffusion/ https://github.com/lstein/stable-diffusion/
- jw1224 4y agoThe fork linked by OP is MPS-based, I can see GPU usage way up in Activity Monitor. Seems performance doubled after a reboot though :)
- geerlingguy 4y agoWeird, on M1 Max Mac Studio, only getting 1.42 it/s :/
- nathas 4y agoI got my units backwards :sweat: My bad!
- black3r 4y ago
- blagie 4y agoHow large an image will this handle (versus how much RAM you have)? It seems the GPU memory requirements beyond 512x512 are obscene.
- michaelchisari 4y agoI'm on a iMac M1 16gb and I can handle up to 768x768 but since it's shared memory I close out every other application and run things overnight. The biggest issue with apple chips is that the --seed setting doesn't work. I should be able to set a seed to, for instance, 1083958 and if I re-run a command at the same resolution with that seed, I should get the same image every time. This would allow me to test different steps so I could generate a 100 images at 16 steps (which is quite fast) and pick the ones that are most promising and re-render at 64 or 128 steps. But currently you can't do that on apple hardware because of an open issue in PyTorch. Genuinely hoping a fix comes soon, until it is this is more of a novelty than a tool on Apple hardware.
- fragmede 4y agoThere's a partial fix for the seed issue on Reddit.
- michaelchisari 4y agoI can't seem to find it, do you have a link?
- fragmede 4y agoThe seed stuff got deleted off of https://www.reddit.com/r/StableDiffusion/comments/x3yf9i/stable_diffusion_and_m1_chips_chapter_2/ https://www.reddit.com/r/StableDiffusion/comments/x3yf9i/sta... but the changes are https://github.com/CompVis/stable-diffusion/compare/1b3c7acce3a9...fragmede:stable-diffusion:mps_consistent_seed https://github.com/CompVis/stable-diffusion/compare/1b3c7acc....
- pugio 4y ago
- sxp 4y agoIs there a good set of benchmarks available for Stable Diffusion? I was able to run a custom Stable Diffusion build on a GCE A100 instance (~$1/hour) at around 1Mpix per 10 seconds. I.e, I could create a 512x512 image in 2.5 seconds with some batching optimizations. A consumer GPU like a 3090 runs at ~1Mpix per 20 seconds. I'm wondering what the price floor of stock art will be when someone can use https://lexica.art/ https://lexica.art/ as a starting point, generate variations of a prompt locally, and then spend a few minutes sifting through the results. It should be possible to get most stock art or concept art at a price of <$1 per image.
- skybrian 4y agoSo you’re estimating over a thousand generated images an hour and less than a tenth of a cent per image using the A100. If that turns out to be accurate, it seems like some online image generation will included in the price of the stock art. (DreamStudio is charging a bit over one cent per generated image at default settings, depending on exchange rates.)
- sowbug 4y agoRelated: I wrote up instructions for running Stable Diffusion on GCE. I used a Tesla T4, which is probably the cheapest that can handle the original code. If you're spinning up an instance to play with, rather than to batch-process, then cheaper makes more sense because most of the machine's time is spent waiting for you to type stuff and look at the results. https://sowbug.com/posts/stable-diffusion-on-google-cloud/ https://sowbug.com/posts/stable-diffusion-on-google-cloud/
- fleddr 4y agoIt can be even cheaper. Midjourney, in case you appreciate their output, has an unlimited plan for 30$ a month. The only limitation is that if you're an extremely heavy user, they may "relax" you, which means results come in a bit slower. Note that they've been also experimenting with a --beta parameter which basically means the algorithm uses StableDiffusion's algorithm behind the scenes, or you can use any of 4 versions of MidJourney's more stylistic algorithms. So if you don't want to tinker or don't have a high-end GPU, it's a cheap way to play around. I have StableDiffusion running locally but still prefer MidJourney. I enjoy the stylistic output but it's also a highly social way to generate art. Everybody is doing it in the open. Anyway, the stock art part is a hairy subject. You should assume that you AI image is not copyrighted. Which begs the question why they would pay at all.
- butUhmErm 4y agoBetween this and efforts to add 3D dimension to 2D images, I don’t see much of a future for digital multimedia creator jobs. Even TikTok could be an endless stream of ML models. Fears of a tech dystopia may be overblown; the masses will just shut off their gadgets and live simpler if labor markets implode within the traditional political correct economic system we have. Open source AI is on the verge of upending the software industry and copyright. I dig it.
- johnfn 4y agoFor those as keen as I am to try this out, I ran these steps, only to run into an error during the pip install phase: > ERROR: Failed building wheel for onnx I was able to resolve it by doing this: > brew install protobuf Then I ran pip install again, and it worked!
- geerlingguy 4y agoIn the troubleshooting section it mentions running: brew install Cmake protobuf rust To fix onnx build errors. I had the same issue.
- jonplackett 4y agoWhat kind of speed does this run at? Eg. How long to make a 512x512 image at standard settings?
- jw1224 4y agoOn my M1 Pro MBP with 16GB RAM, it takes ~3 minutes.
- pwinnski 4y agoI haven't installed from this link specifically, but I used one of the branches on which this is based a few days ago, so the results should be similar. On a first-gen M1 Mac mini with 8GB RAM, it takes 70-90 minutes for each image. Still feels like magic, but old-school magic.
- mark_l_watson 4y agoThanks for writing this up!! I enjoyed getting TensorFlow running with the M1, although a multi-headed model I was working on wouldn’t run. I just made my Dad’s 101 year old birthday card using OpenAI’s image generating service (he loved it) and when I get home from travel I will use your instructions in the linked article. Any advice for running Stable Diffusion locally vs. Colab Pro or Pro+? My M1 MacBook Pro only has 8G ram (I didn’t want to wait a month for a 16G model). Is that enough? I have a 1080 with 10G graphics memory. Is that sufficient?
- Razengan 4y ago101 years! Congratulations!! Does he own a suspiciously plain gold ring by any chance?
- mark_l_watson 4y agono :-)
- ErneX 4y agoFrom the comments here 8GB is not enough, it will swap a lot and take way more time than a 16GB MacBook.
- r3trohack3r 4y agoI've been playing with Stable Diffusion a lot the past few days on a Dell R620 CPU (24 cores, 96 GB of RAM). With a little fiddling (not knowing any python or anything about machine learning) I was able to get img2img.py working by simply comparing that script to the txt2img.py CPU patch. Was only a few lines of tweaking. img2img takes ~2 minutes to generate an image with 1 sample and 50 iterations, txt2img takes about 10 minutes for 1 sample and 50 generations. The real bummer is that I can only get ddim and plms to run using a CPU. All of the other diffusions crash and burn. ddim and plms don't seem to do a great job of converging for hyper-realistic scenes involving humans. I've seen other algorithms "shape up" after 10 or so iterations from explorations people do online - where increasing the step count just gives you a higher fidelity and/or more realistic image. With ddim/plms on a CPU, every step seems to give me a wildly different image. You wouldn't know that steps 10 and steps 15 came from the same seed/sample they change so much. I'm not sure if this is just because I'm running it on a CPU or if ddim and plms are just inferior to the other diffusion models - but I've mostly given up on generating anything worthwhile until I can get my hands on an nvida GPU and experiment more with faster turn arounds.
- squeaky-clean 4y ago> You wouldn't know that steps 10 and steps 15 came from the same seed/sample they change so much. I don't think this is CPU specific, this happens at these very low number of samples, even on the GPU. Most guides recommend starting with 45 steps as a useful minimum for quickly trialing prompt and setting changes, and then increasing that number once you've found values you like for your prompt and other parameters. I've also noticed another big change sometimes happens between 70-90 steps. It's not all the time and it doesn't drastically change your image, but orientations may get rotated, colors will change, the background may change completely. > img2img takes ~2 minutes to generate an image with 1 sample and 50 iterations If you check the console logs you'll notice img2img doesn't actually run the real number of steps. It's number of steps multiplied by the Denoising Strength factor. So with a denoising strength of 0.5 and 50 steps, you're actually running 25 steps. Later edit: Oh and if you do end up liking an image from step 10 or whatever, but iterating further completely changes the image, one thing you can do is save your output at 10 steps, and use that as your base image for the img2img script to do further work.
- yoyohello13 4y agoHas anybody had success getting newer AMD cards working? ROCm support seems spotty at best, I have a 5700xt and I haven't had much luck getting it working.
- switchers 4y ago6600XT reporting in. Spent a few hours on Windows and WSL2 setup attempts, got no where. I don't run Ubuntu at home and don't want to dual boot just for this. From looking around I think I'd have a better chance on native Ubuntu.
- my123 4y agoBuy an NVIDIA card. ROCm isn't supported in any way on WSL2, but CUDA is. AMD just doesn't invest in their developer ecosystem. Also as you use a 6600 XT, no official ROCm support for the die that you use. Only for navi21.
- MintsJohn 4y agoOr wait, if its just about stable diffusion multiple people try to create onnx and directml forks of the models/scripts, which atleast in theory can work for AMD gpus in windows and wsl2
- geerlingguy 4y agoI've tried using this set of steps [1], but have so far not had luck, mostly because the ROCm driver setup is throwing me for a loop. Tried it with an RX 6700 XT and first was going to test on Ubuntu 22.04 but realized ROCm doesn't support that OS yet, so tried again on 20.04 and ended up breaking my GPU driver! [1] https://gist.github.com/geerlingguy/ff3c3cbcf4416be2c0c1e0f836a8183d https://gist.github.com/geerlingguy/ff3c3cbcf4416be2c0c1e0f8...
- my123 4y agoYes. That's expected. AMD market segmented their RDNA2 support in ROCm to the Navi21 set only (6800/6800 XT/6900 XT). It is not officially supported in any way on other RDNA2 GPUs. (Or even on the desktop RDNA2 range at all, that only works because their top end Pro cards share the same die)
- gonehome 4y agoThanks for this - it's rare to see a setup guide that actually works on each step! I did need to run the troubleshooting step too, could probably just move that up as a required step in the guide.
- bfirsh 4y agoIt isn't required for some (most?) users. Weirdly sometimes pip is picking up the wheel for `onnx`, sometimes it isn't, and we can't figure out why. Any Python packaging experts know what's going on? all macOS 12, arm64, Python 3.10. Can't think it wouldn't resolve the wheel. But yes, good idea to move up. I'll stick it next to the `pip install`.
- yboris 4y agoPlease consider also adding a small note to help those few that get stuck with this bug: RuntimeError: expected scalar type BFloat16 but found Float The solution is easy: append the execution command with `--precision full`
- qabqabaca 4y agoCommenting for visibility, thanks for this
- forrestwilkins 4y agoThis worked for me as well, thanks!
- BalogunofAfrica 4y agoCommenting for more visibility, this worked for me too. Thank you!
- orf 4y agoDifferent versions of pip.
- jw1224 4y agoAre we being pranked? I just followed the steps but the image output from my prompt is just a single frame of Rick Astley... EDIT: It was a false-positive (honest!) on the NSFW filter. To disable it, edit txt2img.py around line 325. Comment this line out: x_checked_image, has_nsfw_concept = check_safety(x_samples_ddim) And replace it with: x_checked_image = x_samples_ddim
- pja 4y agoThat means the NSFW filter kicked in IIRC from reading the code. Change your prompt, or remove the filter from the code.
- johnfn 4y agoHaha, busted!
- pja 4y agoTo be fair, the reason the filter is there is that if you ask for a picture of a woman, stable diffusion is pretty likely to generate a naked one! If you tweak the prompt to explicitly mention clothing, you should be OK though.
- fprog 4y agoWow, is that true? I’ve never heard a more textbook ethical problem with a model.
- sp332 4y agoAny chance of this running on an M1 iPad Pro?
- djhworld 4y agoNote that once you run the python script for the first time it seems to download a further ~2GB of data
- nonethewiser 4y agoIncluding a rick astley image for the first thing you gen -_-
- ErneX 4y agoThat’s the NSFW filter :D
- johnfn 4y agoHm, when I run the example, I get this error: > expected scalar type BFloat16 but found Float Has anyone seen this error? It's pretty hard to google for.
- nathas 4y agoYeah. Try running with PYTORCH_ENABLE_MPS_FALLBACK=1 <script> --full-precision
- bfirsh 4y agoAre you running macOS >=12.3?
- johnfn 4y agoOh, no I'm not, I'm on 12.0. Does this make a difference?
- johnfn 4y agoUpdate: I solved this error more properly by upgrading to the latest version. Thanks bfirsh.
- wuyishan 4y agoI am having the same issue on MacOS 12.2.1 (21D62); Python 3.10.6 What did you upgrade to solve this? Thanks! (I can get it working with `--precision full`)
- 4y ago
- TekMol 4y agoDoes running it locally give you anything over using the web version?
- bfirsh 4y agoYou can hack on it, modify it, integrate it with other code, etc!
- johnfn 4y agoAlso, of course, it's entirely free. The web version is actually paid, though it's hard to tell because they're not super transparent about the fact that you're steadily eating through a quota of initial tokens.
- schleck8 4y agoYou can finetune the model with Textual Inversion. And you don't have the safety mechanism for nudity.
- Gigachad 4y agoThe online ones are usually bogged down with too much traffic. Locally seems to be much easier/faster
- deleted 4y ago[deleted]
- usehackernews 4y agoMagnusviri[0], the original author of the SD M1 repo credited in this article, has merged his fork into the Lstein Stable Diffusion fork. You can now run the Lstein fork[1] with M1 as of a few hours ago. This adds a ton of functionality - GUI, Upscaling & Facial improvements, weighted subprompts etc. This has been a big undertaking over the last few days, and I highly recommend checking it out. See the mac m1 readme [3] [0] https://github.com/magnusviri/stable-diffusion https://github.com/magnusviri/stable-diffusion [1] https://github.com/lstein/stable-diffusion https://github.com/lstein/stable-diffusion [2] https://github.com/lstein/stable-diffusion/blob/main/README-Mac-MPS.md https://github.com/lstein/stable-diffusion/blob/main/README-...
- solarkraft 4y agoI ran into: ImportError: cannot import name 'TypeAlias' from 'typing' (/opt/homebrew/Caskroom/miniconda/base/envs/ldm/lib/python3.9/typing.py)
- icedchai 4y agoI ran into this. You need Python 3.10. I had to edit environment-mac.yaml and set python==3.10.6 ...
- dgreensp 4y agoI'm working on getting this running. Instead of "venv/bin/activate" I had to run "source venv/bin/activate". And I got an error installing the requirements, fixed by running "pip install pyyaml" as a separate command.
- valley_guy_12 4y agoHaving to use "source" means you have an older version of conda. Python package management is kind of a mess.
- dgreensp 4y agoWow, I'm getting as low as 1.2 seconds per "step" (about a minute for a 512x512 image with default settings) on my 32 GB M1 laptop (2021, 16-inch).
- bfirsh 4y agoThere's a little "." before "venv/bin/activate" that's easy to miss. I'll update it to "source" to make it more obvious.
- dgreensp 4y agoGreat. FYI, the issue with pyyaml was caused by the "--pre" flag. The reason installing it separately fixes the problem is because it installs it without the "--pre" flag, and then it is already installed when you install the requirements. I haven't been able to make sense of n-samples and n-iters. Changing the former caused generation to freeze at 0%. Changing the latter seems to generate multiple images, even though the SD docs for n-iters are "sample this often".
- rhacker 4y agoI wonder if this is going to be a huge boon to m1 sales.
- sgt101 4y agomight be easier to wait for Diffusers to merge the pull request...
- sroussey 4y agoThis should be put into a docket image to avoid various potential conflicts with locally installed libraries. Anyone do this for the M1?
- bfirsh 4y agoUnfortunately this can't run in Docker because Docker for Mac can't access the M1 GPU. (Several layers of virtualization and emulation!)
- sroussey 4y agoAh, thanks. Yes, that makes sense.
- schleck8 4y agoConda environment
- hnarayanan 4y agoDo you want to add more layers to make it extra slow?
- sgt101 4y agoalso brew upgrade not brew update
- ChildOfChaos 4y agoIs there anyway to keep up with this stuff / beginners guide? I really want to play around with it but it's kinda confusing to me. I don't have an M1 Mac, I have an Intel one with an AMD GPU, not sure if i can run it? don't mind if it's a bit slow, or what is the best way of running it in the cloud? Anything that can product high res for free?
- yreg 4y agoHave you managed to set it up? I might have the same computer as you.
- ChildOfChaos 4y agoNot yet, I haven't had much time to look into it all yet. Looks like it's going to be a lot of fun though.
- Karuma 4y agoYes, you can run it on your Intel CPU: https://github.com/bes-dev/stable_diffusion.openvino https://github.com/bes-dev/stable_diffusion.openvino And this should work on an AMD GPU (I haven't tried it, I only have NVIDIA): https://github.com/AshleyYakeley/stable-diffusion-rocm https://github.com/AshleyYakeley/stable-diffusion-rocm There are also many ways to run it in the cloud (and even more coming every hour!) I think this one is the most popular: https://colab.research.google.com/github/altryne/sd-webui-colab/blob/main/Stable_Diffusion_WebUi_Altryne.ipynb https://colab.research.google.com/github/altryne/sd-webui-co...
- holoduke 4y agofollow this guide: https://github.com/lstein/stable-diffusion/blob/main/README-Mac-MPS.md https://github.com/lstein/stable-diffusion/blob/main/README-... i am runnig it on my 2019 intel macbook pro. 10 minutes per picture
- EddySchauHai 4y agohttps://beta.dreamstudio.ai/dream https://beta.dreamstudio.ai/dream It's not free but I've played with it a lot over the last two days for around $10, generating the most complex photos I can (1024x1024, 150 steps, 9 images, etc)
- code51 4y agoWithout k-diffusion support, I don't think this replicates Stable Diffusion experience: https://github.com/crowsonkb/k-diffusion https://github.com/crowsonkb/k-diffusion Yes, running on M1/M2 (MPS device) was possible with modifications. img2img and inpainting also works. However you'll run into problems when you want k-diffusion sampling or textual inversion support.
- Birch-san 4y agostable-diffusion supports k-diffusion just fine on M1. You just have to detach a tensor in to_d() to stop the values exploding to infinity. https://twitter.com/Birchlabs/status/1563622002581184517?s=20&t=kt9noe6QjD2WLvcQGKbnww https://twitter.com/Birchlabs/status/1563622002581184517?s=2...
- code51 4y agoI've been following your MPS branch and have run it but couldn't address the issue without this explanation. Thank you!
- jclardy 4y agoIs there a proper term to encapsulate M1/M2 Macs now that we have the M2? IE Apple Silicon Macs works but is a bit long. MX Macs? M-Series? ARM Macs?
- gzer0 4y agoThe difference between an M2 air (8gb/512gb) versus an M1 pro (16gb/1tb) is much more than I expected. * M1 pro (16gb/1tb) can run the model in around 3 minutes. * M2 air (8gb/512gb) takes ~60 minutes for the same model. I knew there would be some throttling due to the m2 air's fanless model, but I had no idea it would be a 20x difference (albeit, the m1 pro does have double the RAM. I don't have any other macbooks to test this on).
- qayxc 4y agoI suspect the lack of RAM is the issue here.
- andybak 4y agoUnscientifically that puts the M1 Pro GPU at about 25% of the performance of a RTX 3080. Not too shabby... EDIT - this comment implies it's much faster: https://news.ycombinator.com/item?id=32679518 https://news.ycombinator.com/item?id=32679518 If that's correct then it's close to matching my 3080 (mobile).
- fassssst 4y agoimg2img runs in 6 seconds on my GeForce 3080 12 GB. 6+ it\s depending on how much GPU memory is available. If I have any electron apps running it slows down dramatically.
- fragmede 4y agowhat args are you passing to img2img?
- andybak 4y agoCurious about: 1. Image size 2. Steps 3. What your numbers are for text2img 4. (most importantly) are you including the 30 seconds or so it takes to load the model initially? i.e. if you were to run 10 prompts and then divide the total time by 10, what are your numbers?
- moneycantbuy 4y agoAnyone know the largest possible image size > 512x512? I'm getting the following error when trying 1024x1024 with 64 GB RAM on M1 MAX: /opt/homebrew/Cellar/python@3.10/3.10.6_2/Frameworks/Python.framework/Versions/3.10/lib/python3.10/multiprocessing/resource_tracker.py:224: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d '
- capableweb 4y agojust 512x require something like 10GB VRAM on your GPU, 1024x would need even more. How much VRAM does the M1 Max GPU have? You're probably running out of memory.
- tgtweak 4y agoM architecture is unified memory - so the system memory is shared as GPU memory. There is likely a cap somewhere on how much can be used by a single app (or collectively by all apps).
- enduser 4y agoI have the same problem with anything over 512x512 on my M1 Ultra with 128GB. VRAM must be capped.
- moneycantbuy 4y agoThanks for the Ultra data point. I'm able to get 768x896 to run, but the output image is still white noisy at 50 ddim steps, perhaps related to the phenomena of being trained/windowed on 512x512 images as sibling squeaky-clean described. RAM usage at various sizes: 512x512 14 GB, 768x768 26 GB, 768x896 32 GB
- squeaky-clean 4y agoThose values seem really high compared to my setup, windows/nvidia/lstein repo. For me 512x512 uses 6.1GB. Random guess but I think your pipeline is running with full-precision floats (32bit), while by default the repo should be using autocast() which will try to use half-precision floats wherever possible. I know an optimizedSD repo exists and one of the steps they take is explicitly setting precision to half. (And other changes that reduce memory usage but decrease iteration speed). However I don't know how M1/Metal handles half-precision, hopefully it doesn't just cast them back to 32bit. Also white noisy images at 50 steps seems off to me. At 50 steps in a large image I definitely get a visible product. It's just often non-euclidean or very scattered bits of organization and chaos.
- deckeraa 4y agoVery nice to see this available for hardware I own. Now I can achieve my dream of a Corporate Memphis + Hieronymus Bosch mashup.
- bamurphymac1 4y agoLow entropy observer A38B gasps, spills his coffee and slams the panic button, triggering the 3758th reboot of the simulation. “Sorry all. We can’t spare the hardware for THAT kind of recursive self improvement. Setting this instance back to 1970. GG”
- adamj9431 4y agoHow is Stable Diffusion on DreamStudio.ai so much faster than the reports here? Seems to only take 5-10 seconds to generate an image with the default settings. I.e. How are they providing access to GPU compute several orders of magnitude more powerful than an M1, for free?
- tgtweak 4y agoA100 devices in the cloud on preemptive/spot instances?
- schleck8 4y ago1. Dreamstudio is paid, with 2 EUR worth of credit for free. 2. The M1 GPU is an iGPU. A good iGPU, sure, but it's not anywhere near the performance of a dedicated, cooled GPU with dedicated VRAM. On a 2060 Super with 8 GB of VRAM and with tensor cores it takes 15 seconds to infer with the default settings. If Dreamstudio uses deep learning GPUs then there is your answer to why it is as fast.
- fleddr 4y agoThe guy behind it is an ex hedge fund manager. Using private funds he's built a massive fleet of A1000s at AWS. So it's an enormous amount of compute created from private funds that he considers to be "for humanity". Currently, he funds it and is also the "GPU overlord", he exclusively decides which applications gets to use it. His plan, or at least his claim, is to transform this situation in it being more diversely funded (institutions, businesses, even the UN) and for access to be decided by committee with main criteria it being useful for humanity. Let's see if he sticks to his word, but I find it inspirational. AI was on a trajectory to be solely in the hands of a hand full of ultra rich companies that can afford to train and run it, and us poor mortals being at the whims of gatekeeper terms. This guy is on a trajectory to put AI in the hands of the people. Not just for art, for everything. If he fully sees this through, he's destined to be a tech icon.
- preommr 4y agoThe guy's name is Emad Mostaque btw There's a recent video interview he did that goes into his vision: https://www.youtube.com/watch?v=YQ2QtKcK2dA https://www.youtube.com/watch?v=YQ2QtKcK2dA
- joshstrange 4y agoIt's insane to me how fast this is moving. I jumped through a bunch of hoops 2-3 days ago to get this running on my M1 Mac's GPU and now it's way easier. I imagine we will have a nice GUI (I'm aware of the web-ui, I haven't set it up yet) packaged in an mac .app by the end of next week. Really cool stuff.
- addandsubtract 4y agoI hope this kickstarts some kind of M1 migration. There are so many ML projects I'd like to try, but they all depend on CUDA.
- joshstrange 4y agoYep, I was just thinking the same thing. M1/M2 appears to be a huge untapped resource for ML stuff as this proves. I maxed out my MBP Max and this is probably the first time I'm actually fully using the GPU cores and it's pretty freaking cool. Creating landscapes or fictional characters (think D&D) is already super fun, I look forward to playing with img2img some more as well.
- zone411 4y agoThe performance gap to the top-end Nvidia cards will get much larger as they release new cards later this year, though.
- jdminhbg 4y agoMaybe, but I can buy a Mac, you just order one from Apple.
- zone411 4y agoIt's not clear if the shortages will happen with this new release as they did last time. Ethereum mining is going away and not as many people are stuck at home because of Covid. On the other hand, the performance increase looks to be substantial, increasing the demand.
- simonebrunozzi 4y agoI don't want to sound lazy, but I would be expecting a .dmg for Macs, and I don't seem to find it. Am I blind, or it simply hasn't been prepared yet?
- omarelbie 4y agoAt the rate it's moving it doesn't seem too far off, but I think it's just a tad too early.
- adrianvoica 4y agoTried "transparent dog", got rickrolled. Why is this NSFW? ...anyway, I disabled the filter and... it's pretty neat! Calling all AI Overlords, soon. :))
- wenbin 4y agoThanks for the writeup! It works smoothly on my M1 Macbook Pro! A few days ago, I tried Stable Diffusion code and was not able to get it work :( Then I gave up... Today, following steps in this blog post, it works for the very first try. Happy!
- imtemplain 4y agoI'm ready to pay for a Windows + AMD GPU guide at this point, why is there no single blogpost on this, please help.
- bobthebl0b 4y agoI was like you, but man $2,000 no thanks. This saved my budget and time https://github.com/glonlas/Stable-Diffusion-Apple-Silicon-M1-Install https://github.com/glonlas/Stable-Diffusion-Apple-Silicon-M1...
- Daegalus 4y agoI'm sure you can do it by installing the ROCm drivers, install pytorch-rocm based on the download page for pytorch, and make sure to set any settings in environment variables. you can try following my Linux AMD guide and see what you can pick and choose to get it working https://yulian.kuncheff.com/stable-diffusion-fedora-amd/ https://yulian.kuncheff.com/stable-diffusion-fedora-amd/
- vvanirudh 4y agoRunning into this error `RuntimeError: expected scalar type BFloat16 but found Float` when I run `txt2img.py`
- yboris 4y agoConfirming I'm stuck on the same error when running the tutorial-instructed python scripts/txt2img.py command RuntimeError: expected scalar type BFloat16 but found Float
- ml_basics 4y agoYes, me too! Please post here if you find a solution for all the other people that come and find this by commmand-F'ing this error
- sytse 4y agoI'm stuck on 'RuntimeError: expected scalar type BFloat16 but found Float' too. Most relevant links seems https://github.com/CompVis/stable-diffusion/pull/47 https://github.com/CompVis/stable-diffusion/pull/47 but I'm not sure. Please post when there is a solution.
- alvb 4y agoThat might have to do with your Mac OS version. Pre-12.4 Mac OS does not allow the Torch backend to use the M1 GPU, and so the script attempts to use the cpu, but then the cpu does not support half-precision numbers.
- alvb 4y agoYep---that was it in my case. I had the same error but it went away after upgrading to MacOS 12.5. You should actually check if your PyTorch installation can detect the mps backend: `torch.backends.mps.is_available()` must be equal to True.
- 4y ago
- moneycantbuy 4y agoWhat's with the ~25% chance of an image being all black? Also, seeds aren't replicating.
- amilios 4y agoHow long does it take to generate a single image? Is it in the 30 min type range or a few mins? It's hypothetically "possible" to run e.g. OPT175B on a consumer GPU via Huggingface Accelerate, but in practice it takes like 30 mins to generate a single token.
- lostmsu 4y agoI was able to run YaLM 100B in about 5min per iteration, NVMe being the bottleneck.
- LanternLight83 4y agoRuns on my 2070S at 12s/image (no batch optimization) and on my GTX1050 4GB at 90s/image
- holoduke 4y agoOn my late 2019 intel macbook pro with 32gb and a AMD 5550m it takes about 7-10 minutes to generate an image.
- Gigachad 4y agoI'm using a 2021 Macbook Pro with the base tier M1 Pro and it generates images in about 1 minutes per image.
- Birch-san 4y agoAbout 10 secs per image on M1 Max with the right noise schedule and sampler. https://twitter.com/Birchlabs/status/1565029734865584143?s=20&t=TnE7e1OhBG-GF1HYZ3GNrQ https://twitter.com/Birchlabs/status/1565029734865584143?s=2...
- mrkstu 4y agoI consistently have items only partially in frame- horses/fish/etc- any tips on getting the algo to keep specified items fully in frame?
- SirYandi 4y agoStruggle with this too. One keyword which helps is 'wide angle'. Sometimes 'full body shot' works for generating humans.
- mdswanson 4y agovirtualenv isn't required. You can just use python -m venv venv and get the same results with one fewer dependency.
- andrethegiant 4y agoYesssss I've been waiting for this!
- _venkatasg 4y agoI keep running into issues, even after installing Rust in my condo environment (using conda). Specifically the issue seems to be building wheels for `tokenizers`: warning: build failed, waiting for other jobs to finish... error: build failed error: `cargo rustc --lib --message-format=json-render-diagnostics --manifest-path Cargo.toml --release -v --features pyo3/extension-module -- --crate-type cdylib -C 'link-args=-undefined dynamic_lookup -Wl,-install_name,@rpath/tokenizers.cpython-310-darwin.so'` failed with code 101 [end of output] note: This error originates from a subprocess, and is likely not a problem with pip. ERROR: Failed building wheel for tokenizers Failed to build tokenizers ERROR: Could not build wheels for tokenizers, which is required to install pyproject.toml-based projects Any suggestions?
- benhalllondon 4y agoI played around a bit and found out dropping the tokenisers version to 0.11.6 worked `pip install tokenizers==0.11.6` first
- _venkatasg 4y agoThat worked thank you!
- msoad 4y agoPlease someone package all of this and the WebUI into an Electron app so common people can also hack on it!
- andrewmunsell 4y agoThere are several packages that provide web UIs, like this one for example: https://github.com/hlky/stable-diffusion-webui https://github.com/hlky/stable-diffusion-webui It's not quite the ease of setup of an Electron app, but once setup it's pretty easy to use.
- gregsadetsky 4y agoBananas. Thanks so much... to everyone involved. It works. 14 seconds to generate an image on an M1 Max with the given instructions (`--n_samples 1 --n_iter 1`) Also, interesting/curious small note: images generated with this script are "invisibly watermarked" i.e. steganographied! See https://github.com/bfirsh/stable-diffusion/blob/main/scripts/txt2img.py#L253 https://github.com/bfirsh/stable-diffusion/blob/main/scripts...
- grishka 4y ago> Also, interesting/curious small note: images generated with this script are "invisibly watermarked" i.e. steganographied! Why?
- Retr0id 4y agoSo that future iterations of StableDiffusion (or similar models) don't end up getting trained on their own outputs.
- hifikuno 4y agoOh wow, I didn't even think of that. I am pretty sure a few of the repo's have turned off the invisible watermark, I wonder if that will have consequences down the line for training data.
- gregsadetsky 4y ago... so this means that watermarking an image you own is probably the only way to avoid it being used for training further models? :-)
- cageface 4y agoAfter playing around with all of these ML image generators I've found myself surprisingly disenchanted. The tech is extremely impressive but I think it's just human psychology that when you have an unlimited supply of something you tend to value each instance of it less. Turns out I don't really want thousands of good images. I want a handful of excellent ones.
- chrisfrantz 4y agoHuman curation will likely remain valuable into the future.
- Gigachad 4y agoWhat's this log message about when generating an image? Creating invisible watermark encoder (see https://github.com/ShieldMnt/invisible-watermark https://github.com/ShieldMnt/invisible-watermark)...
- gregsadetsky 4y agoSee my sibling comment here https://news.ycombinator.com/item?id=32684003 https://news.ycombinator.com/item?id=32684003 :-) The code that generates it is here: https://github.com/bfirsh/stable-diffusion/blob/main/scripts/txt2img.py#L253 https://github.com/bfirsh/stable-diffusion/blob/main/scripts... You can remove `img = put_watermark(img, wm_encoder)` which appears at lines 317 and 333 to get rid of the watermarking.
- js2 4y agoA few suggested changes to the instructions: /opt/homebrew/bin/python3 -m venv venv # [1, 2] venv/bin/python -m pip install -r requirements.txt # [3] venv/bin/python scripts/txt2img.py ... 1. Using /opt/homebrew/bin/python3 allows you to remove the suggestion about "You might need to reopen your console to make it work" and ensures folks are using the just installed via homebrew python3, as opposed to Apple's /usr/bin/python3 which is currently 3.8. It also works regardless of the user's PATH. We can be fairly confident /opt/homebrew/bin is correct since that's the standard homebrew location on Apple Silicon and folks who've installed it elsewhere will likely know how to modify the instructions. 2. No need to install virtualenv since Python 3.6 which ships with a built-in venv module which covers most use cases. 3. No need to source an activate script. Call the python inside the virtual environment and it will use the virtual environment's packages.
- schappim 4y agoI just got rick-rolled by the model. Using the prompt: "1990s textbook background mephis style"[sic] (yup I meant memphis)[0], I got back this: [1]. Rerunning the same prompt, I got: [2]. [0] https://files.littlebird.com.au/Shared-Image-2022-09-02-10-29-53-RILg9J.png https://files.littlebird.com.au/Shared-Image-2022-09-02-10-2... [1] https://files.littlebird.com.au/grid-0004-2xXAGF.png https://files.littlebird.com.au/grid-0004-2xXAGF.png [2] https://files.littlebird.com.au/grid-0005-kcfgq7.png https://files.littlebird.com.au/grid-0005-kcfgq7.png
- deleted 4y ago[deleted]
- fastball 4y agoThe rick-roll is the NSFW filter. It's not actually that great at detecting NSFW content.
- schappim 4y agoThanks! You of course are right :) txt2img.py line 324 update to x_checked_image = x_samples_ddim #, has_nsfw_concept = check_safety(x_samples_ddim) ... disables the NSFW so you don’t get the rickroll images. It seems to be over enthusiastic in terms of its NSFW detection
- bschwindHN 4y agoEveryone posting their pip/build/runtime errors is everything that's wrong with tooling built on top of python and its ecosystem. It would be nice to see the ML community move on to something that's actually easily reproducible and buildable without "oh install this version of conda", "run pip install for this package", "edit this line in this python script".
- smoldesu 4y agoI found a pretty great Docker image for the SD webui. I was forced to extreme measures since NixOS isn't super friendly with Conda (and I was particularly lazy). Worked out fine in the end, though. Highly recommended if you're on an Nvidia rig: https://github.com/AbdBarho/stable-diffusion-webui-docker https://github.com/AbdBarho/stable-diffusion-webui-docker
- sanroot99 4y agoDocker the silver bullet /s
- nl 4y agoI don't think this is particularly fair. This is literally hours old, and people installing now are really debugging rather than installing a "finished" build. The forks are weird mashups of bits of repos, and running on a M1 GPU is something that barely works itself. Give it maybe 3 months and it will be much smoother.
- bschwindHN 4y agoI think it is fair, actually. This is not unique to hours-old python projects, it's a common theme among almost all python tools I've used. I have a suspicion that that something written in Julia, Go, Rust, or possibly even C wouldn't have nearly this many issues. I'm not talking about debugging the actual functionality of the software, but rather the environment and tooling surrounding the language and software built with it. This project in particular should be an easy case because you know the hardware you'll be running on ahead of time. I'm ranting a bit, but I've tried so many tools based on python and almost _none_ of them built/installed/ran correctly on the happy path laid out in each project's readme. Anyway, sorry, rant over.
- avocado2 4y ago
- e40 4y agoFor me: File "/Users/layer/src/stable-diffusion/venv/lib/python3.10/site-packages/torch/serialization.py", line 250, in __init__ super(_open_file, self).__init__(open(name, mode)) FileNotFoundError: [Errno 2] No such file or directory: 'models/ldm/stable-diffusion-v1/model.ckpt' The directory is empty. Hmm. I forgot to mv sd-v1-4.ckpt models/ldm/stable-diffusion-v1/model.ckpt On a Mac Studio data: 100%|| 1/1 [00:43<00:00, 43.20s/it] Sampling: 100%|| 1/1 [00:43<00:00, 43.20s/it]
- deleted 4y ago[deleted]
- Yido 4y agoInteresting!
- RosanaAnaDana 4y agoThis whole 2 month period has felt like the first few steps onto some kind of exponential.
- e40 4y agoI just found out that Activity Monitor doesn't show GPU activity. :(
- Linda703 4y ago[dead]
- dzink 4y agoIf you have a top of the line M1 MBP but the hard drive is 2TB, would it make sense to plug in an external hard drive for the 4TB model or would it render the effort futile due to performance issues?
- ggerganov 4y agoThe model is 4GB, not 4TB
- deleted 4y ago[deleted]
- Myrmornis 4y agoThe various articles/tutorials seem a bit confusing: even though they say "M1", they also worked fine for me on an Intel Mac (and does end up using GPU). Does anyone know how to think about the --W --H and --f flags to create larger images? I have 64GB memory, but I get errors from PyTorch saying things like "Invalid buffer size: 7.54 GB" when I try to increase W and H, and I haven't managed to make the Python process use more than about 15GB by playing around so far.
- wut42 4y agoThat must be using CUDA then, and you need a gpu with at least 8GB of VRAM, afaik (not RAM).
- Myrmornis 4y agoAh-ha, thanks. My Intel Mac (2019) has https://www.techcenturion.com/intel-uhd-graphics-630 https://www.techcenturion.com/intel-uhd-graphics-630 > Being an Integrated GPU, the Intel UHD Graphics 630 doesn’t have any Video/Graphics Memory of its own. Instead, it utilizes the system’s memory (RAM) dynamically for the same purpose. You can change the maximum Video Memory from the BIOS settings. I get the impression Apple isn't going to give me much control over that (https://www.reddit.com/r/macmini/comments/e82knm/can_i_set_how_much_ram_is_available_for/ https://www.reddit.com/r/macmini/comments/e82knm/can_i_set_h...). I did have "Automatic Graphics Switching" enabled, but I'm seeing much the same with it off.
- deleted 4y ago[deleted]
- chromejs10 4y agoI keep getting `No module named 'ldm'` after I run `python scripts/dream.py --full_precision`. I've confirmed 'ldm' is activate in conda. Any idea?
- abc_lisper 4y agoI get the same thing too!
- vesterde 4y agoTry to put `sys.path.insert(0, os.getcwd())` after `import sys` in dream.py. That fixed it for me.
- totetsu 4y agoWatching pythons progress bar and waiting this long for an image, I feel like we've come full circle to the days of dail-up modems
- dzink 4y agoIf you have an M1 MAX with 64GB of memory you can make your images bigger. 512x512 only takes 13GB :)
- nicolashahn 4y agohow do you do this?
- shagie 4y agoAs a side bit, this model appears to have difficulty with the prompt "wolf with bling walking down a street" and often generates an image that I am fairly sure is not unique and is not representative of the idea that is trying to be communicated in that text.
- _yb2s 4y agoAnyone else surprised that the results just aren't very good? I followed the instructions and it works, but just seems kinda-like deep dream level results circa 2005. Lots of blurry eyes in the wrong spots, most objects seem really blurrly and cut off. Nothing like the demo examples I've seen online. Are the default options not optimal to get the best results?
- bobthebl0b 4y agoThanks for this tutorial, I had errors and spent time to fix them and I found this script that install LStein project on M1: https://github.com/glonlas/Stable-Diffusion-Apple-Silicon-M1-Install https://github.com/glonlas/Stable-Diffusion-Apple-Silicon-M1... On my side this helped me to make it works. I ran it and it was installed.
- mikhael28 4y agoVery interestingly, this is the first true use case I have noticed where new, bleeding edge technology is much better, seemingly, on M1 than Intel GPUs.