16 ms·
Stable Diffusion with Core ML on Apple Silicon
- calrizien 4y agoWhere is the community for this project?
- tosh 4y agoAtila from Apple on the expected performance: > For distilled StableDiffusion 2 which requires 1 to 4 iterations instead of 50, the same M2 device should generate an image in <<1 second https://twitter.com/atiorh/status/1598399408160342039 https://twitter.com/atiorh/status/1598399408160342039
- cammikebrown 4y agoIf you told me this was possible when I bought an M1 Pro less than a year ago, I wouldn’t believe you. This is insane.
- ncr100 4y agoAgreed. And the posted benchmarks for the M2 Macbook Air make me consider 'upgrading' to an Air.
- Terretta 4y agoThat laptop feels like liquid power. It's uncanny. Macbook Airs (way back when) felt sluggish. The MBA M1 changed that, it was "fine". These M2s are unexpectedly responsive on an ongoing basis. The MacBook Pro M1 Max is great (would be fantastic except they lost a Thunderbolt port in favor of legacy HDMI and memory card jacks), but you expect that machine to be responsive, so it's less surprising. The Studio Ultra, though, never slows down for anything. Still, if the Air could drive two external screens instead of one, I'd "downgrade" from the Max.
- jclardy 4y agoI'd give the M1 air more credit - I moved from a 2019 16" Pro to the Air and performance was nearly identical except for long running tasks (> 10 minutes.) So for mobile app builds, it was blazing fast. And in the meantime the intel machine was blaring fans after the first 30 seconds while the Air barely got warm.And then the real kicker was watching the battery on the intel machine visibly dropping a few percentage points, while the air sits at the same level the whole time. I've since moved to the M2 air, and it is noticeably faster than M1, but it isn't the huge leap from last gen intel that the M1 was. But the hardware itself feels way better.
- SXX 4y agoI dont like lack of open source drivers, but honestly for work DisplayLink works just fine on MacOS. E.g I used 4 monitors on M1 Air using DisplayLink: * Air built-in display * 2K display connected via USB-C -> DisplayPort adapter * Two more 2K displays of same model via DisplayLink connected via USB hub For all practical means it's almost impossible to see any DisplayLink compression artifacts even in most of games. PS: Each adapter cost me $40: https://www.amazon.com/gp/product/B08HN2X88P/ https://www.amazon.com/gp/product/B08HN2X88P/
- Terretta 4y agoAppreciate this reply, TY for sharing the exact product that's working for you! Been nervous to dip into it, given the architecture change and last year's challenges with display link docks. // UPDATE: Oops, looking at the product, I see I should have specified: 4K screens or higher. About half our desks are 2 x 4K, about half 2 x 5K, except the Air M1 folks who are 1 x 5K.
- SXX 4y agoSadly I can only report it working on 2560x1440. Even though lower resolution is specified on Amazon. For higher resolution some other solution is required.
- peppertree 4y agoLast nail in the coffin for DALL·E.
- m00dy 4y agoyeah, finally we see the real openAI
- visarga 4y agomore open than open source, it's the open model age
- astrange 4y agoI think they can move upmarket just as well as anyone else.
- mensetmanusman 4y agoNot really, everyone will have their own flavor on how to rapidly train the model. Dall-e et. al will still be able to bandwagon off of all the free ecosystem being built around the $10M SD1.4 model that is showing what is possible. E.g. Dall-e could go straight to Hollywood if their model training works better than SD’s. The toolsets will work
- swyx 4y agosource for the $10m number? i havent heard that one before, everyone just keeps parrotting the 600k single run number that is obviously misleading
- nomel 4y agoThe true metric contains the output quality of the image, not just the speed. DALL-E output is, generally, much better for things that aren't standard looking.
- Terretta 4y agoIf that's the metric, MidJourney --v 4 --q 2 is the leader, and it's not close.
- chasd00 4y agoi'm very ignorant here so forgive me but if it can generate images that fast can it be used to generate a video?
- vletal 4y agoYeah, sure. The issue is with temporal consistency. Meta and Google have some successes in that area. https://mezha.media/en/2022/10/06/google-is-working-on-imagen-video-its-own-video-generating-ai/ https://mezha.media/en/2022/10/06/google-is-working-on-image... Give it some time and SD will be able to do the same.
- gcanyon 4y agoThere are different requirements for generating video -- at a minimum, continuity is tough. There are models for producing video, but (as far as I've seen) they're still a bit wobbly.
- valgaze 4y agoVideo is really a series of frames, the framerate for film/human can get away with 24 frames/second-- so maybe ~40ms/image for real-time at least? What's cool about the era in which we live is if you look at high-performance graphics for games or simulations, for instance, it may in fact be faster to a the model to "enhance" a low-resolution frame rather than trying to render it fully on the machine. ex. AMD's FSR vs NVIDIA DLSS - AMD FSR (Fidelity FX Super Resolution): https://www.amd.com/en/technologies/fidelityfx-super-resolution https://www.amd.com/en/technologies/fidelityfx-super-resolut... - NVIDIA DLSS (Deep Learning Super Sampling): jhttps://www.nvidia.com/en-us/geforce/technologies/dlss/ https://www.nvidia.com/en-us/geforce/technologies/dlss/ AMD's approach renders the game at a crummy, low-detail resolution then each frame uses "upscales" Both FSR and DLSS aim to improve frames-per-second in games by rendering them below your monitor’s native resolution, then upscaling them to make up the difference in sharpness. Currently, FSR uses spatial upscaling, meaning it only applies its upscaling algorithm to one frame at a time. Temporal upscalers, like DLSS, can compare multiple frames at once, to reconstruct a more finely-detailed image that both more closely resembles native res and can better handle motion. DLSS specifically uses the machine learning capabilities of GeForce RTX graphics cards to process all that data in (more or less) real time. Video is really a series of frames, the framerate for film/human could get away with 24 frames/second-- ~40ms/image for real-time. What's cool about the era in which we live is if you look at high-performance graphics for games or simulations, it may in fact be faster to run the model on each frame to "enhance" a low-resolution frame rather than trying to render it fully on the machine. ex. AMD's FSR vs NVIDIA DLSS - AMD FSR (Fidelity FX Super Resolution): https://www.amd.com/en/technologies/fidelityfx-super-resolution https://www.amd.com/en/technologies/fidelityfx-super-resolut... - NVIDIA DLSS (Deep Learning Super Sampling): https://www.nvidia.com/en-us/geforce/technologies/dlss/ https://www.nvidia.com/en-us/geforce/technologies/dlss/ AMD's approach renders the game at a crummy, low-detail resolution then use "spatial upscaling" to enhance the images one frame at a time. NVIDIA DLSS uses "temporal upscaling" to pass over multiple frames and uses other capabilities exclusive to Nvidia's cards to stitch together the frames. This is a different challenge than generating the content from scratch I don't think this is possible in real-time yet, but someone put a filter trained on the German country side to produce photorealistic Grand Theft Auto driving gameplay: https://www.youtube.com/watch?v=P1IcaBn3ej0 https://www.youtube.com/watch?v=P1IcaBn3ej0 Notice the mountains in the background go from Southern California brown to lush green https://www.rockpapershotgun.com/amd-fsr-20-is-a-more-demanding-higher-quality-upscaling-upgrade#:~:text=Currently%2C%20FSR%20uses%20spatial%20upscaling,and%20can%20better%20handle%20motion https://www.rockpapershotgun.com/amd-fsr-20-is-a-more-demand....
- hbn 4y agoSD2 is the one that was neutered, right? Maybe a dumb question but can the old model still be run?
- qclibre22 4y agoAlso, can you not "upgrade" but still run new models?
- astrange 4y agoYou can do anything you want. SD2 wasn’t “neutered”, the piece of it from OpenAI that knew a lot of artist names but wasn’t reproduceable was replaced with a new one from Stability that doesn’t. You can fine-tune anything you want back in.
- l33tman 4y agoThe training-set was nerfed really good as well, it wasn't just OpenCLIP that was replaced. They will successively re-admit more training data during the 2.x releases I guess.
- astrange 4y agoYes, they removed some NSFW which might've hurt it, but releasing models that can generate CP /will/ get you in legal trouble. The "in the style of Greg Rutkowski" prompts from SD1 though, IIRC, were thought to be proof it was reproducing the training set. But it actually only saw ~27 images of his, and the rest was residual biases from CLIP.
- kyleyeats 4y agoIt's less versatile out of the box. Give it a couple months for the community to catch up. Everyone is still figuring out what goes where, and SD 1.x was "everything goes in one spot." It was cool and powerful, but limited.
- minimaxir 4y ago
- mrtksn 4y agoWith the full 50 iterations it appears to be about 30s on M1. They have some benchmarks on the github repo: https://github.com/apple/ml-stable-diffusion https://github.com/apple/ml-stable-diffusion For reference, previously I was getting about <3 minutes for 50 iterations on my Macbook Air M1. I haven't yet tried Apple's implementation but it looks like a huge improvement. It might take it from "possible" to "usable".
- liuliu 4y agoYeah, it is just PyTorch MPS backend is not fully baked and have some slowness. You should be able to get close to that number with maple-diffusion (probably 10% slower) or my app: https://drawthings.ai/ https://drawthings.ai/ (probably around 20% slower, but it supports samplers that takes less steps (50 -> 30)).
- washadjeffmad 4y agoFor comparison, it's also taking ~3min @ 50 iterations on my 12c Threadripper using OpenVino. It sounds like the improvements bring the M1 performance roughly in line with a GTX 1080.
- mrtksn 4y agoI have Macbook Air M1, which is passively cooled. When cooled properly, that is thermal pad mod combined with a fan under the laptop, I'm getting closer to 2min - something like 2.8s per iteration. I guess it would be something 140s for 50 iterations on a MacBook Pro or Mac mini for M1.
- desro 4y agoThis is accurate re: M1 Mac Mini times IME
- joakleaf 4y agoThe Apple Neural Engine in the m1 is supposed to be able to perform 11 tops. The GTX 1080 about 9-11 tflops. So sounds plausible that the m1 can reach the same level in some use cases with the right optimizations.
- minimaxir 4y agoNote that this is extrapolation for the distilled model which isn't released quite yet. (but it will be very exciting when it does!)
- neonate 4y agohttps://github.com/apple/ml-stable-diffusion https://github.com/apple/ml-stable-diffusion
- christiangenco 4y agoOh gosh that's an intimidating installation process. I'll be much more interested when I can just `brew install` a binary.
- MuffinFlavored 4y agoI could be wrong but I think part of the issue is this needs some large files for the trained dataset?
- deleted 4y ago[deleted]
- artimaeis 4y agoA bit different take is DiffusionBee, if you're curious to try it out in a GUI form. https://diffusionbee.com https://diffusionbee.com
- aryamaan 4y agodoes it use the optimised model for Apple chips?
- belthesar 4y agoNot yet, likely, but the project is very active. I could see it coming quite soon.
- Gigachad 4y agoI just tested that app and it was taking about 1s/it using the "Double quality, double time" version. Spat out quite nice images at 25 iterations. Way better than stuff I had tried before which looked worse after a minute than this generates in 25 seconds.
- pkage 4y agoHow does this compare with using the Hugging Face `diffusers` package with MPS acceleration through PyTorch Nightly? I was under the impression that that used CoreML under the hood as well to convert the models so they ran on the Neural Engine.
- liuliu 4y agoIt doesn't. MPS largely is on GPU. PyTorch's MPS implementation is incomplete a few weeks ago as well. This is about 3x faster.
- wincy 4y agoIs it? I just ran it on my M1 MacBook Air and am getting 3 it/sec, same as I was using Stable Diffusion for M1. Maybe I'm doing something wrong?
- liuliu 4y agoThat's surprising to me, although I did the look about 3 weeks ago, and MPS support is a moving target. It is just M1 without Pro or Ultra right? Also, diffusers does support different backends other than PyTorch.
- deleted 4y ago[deleted]
- mark_l_watson 4y agoGreat stuff. I like that they give directions for both Swift and Python This gets you text descriptions to images. I have seen models that given a picture, then generate similar pictures. I want this because while I have many pictures of my grandmothers, I only have a couple of pictures of my grandfathers and it would be nice to generate a few more. Core ML is so well done. A year ago I wrote a book on Swift AI and used Core ML in several examples.
- astrange 4y agoThat’s DreamBooth. There are some services that will do it for you.
- mark_l_watson 4y agoThanks!
- mromanuk 4y agoI’m making one of those services, if you are interested, please reach me at my email. I would like to know what you have in mind regarding your grandmothers
- behnamoh 4y agoThis may sound naive, but what are some use cases of running SD models locally? If the free/cheap options exist (like running SD on powerful servers), then what's the advantage of this new method?
- tosh 4y agoWorks offline, privacy, independent of SaaS (API stability, longevity, …). I'm sure there are more.
- gjsman-1000 4y agoPowerful servers with GPUs are expensive. Laptops you already own, aren't.
- sofaygo 4y ago> There are a number of reasons why on-device deployment of Stable Diffusion in an app is preferable to a server-based approach. First, the privacy of the end user is protected because any data the user provided as input to the model stays on the user's device. Second, after initial download, users don’t require an internet connection to use the model. Finally, locally deploying this model enables developers to reduce or eliminate their server-related costs.
- huggingmouth 4y agoStability! The main reason why I use it locally is because I don't want some random dev unilaterally deciding to change or "sunsetting" features I rely on. Centralized services small and large are guilty of this and I'm sick of it.
- yazaddaruvala 4y ago"Hey Siri, draw me a purple duck" and it all happens without an internet connection! If you mean monetary usecases: Roughly something like Photoshop/Blender/UnrealEngine with ML plugins that are low latency, private, and $0 server hosting costs.
- jwitthuhn 4y ago
- deleted 4y ago[deleted]
- zimpenfish 4y agoMan, this takes a ton of room to do the CoreML conversions - ran out of space doing the unet conversion even though I started with 25GB free. Going on a delete spree to get it up to 50GB free before trying again.
- pyinstallwoes 4y agoHow much space do you have and how much do you try to keep free? I get freaked out if I have less than 400gb free.
- zimpenfish 4y ago/dev/disk3s5 926Gi 857Gi 52Gi 95% 8067489 540828800 1% /System/Volumes/Data It normally hovers around 30-35Gi free.
- password4321 4y agoAll hail Grand Perspective back in the day, not sure who is carrying the "what's wasting my disk space" torch for free these days. Edit: still alive! https://grandperspectiv.sourceforge.net/ https://grandperspectiv.sourceforge.net/
- peddling-brink 4y agoncdu is the best in my book. TUI, supports deletion of files and folders, and very simple to understand. GUI apps for this task like GP and the like are more visually complex than they need to be.
- password4321 4y agoGood point! One gotcha for me is ncdu2 going Zig and Zig dropping support for OS versions as Apple does.
- astrange 4y agoOmniDiskSweeper is a GUI that isn’t complex.
- sorenjan 4y agoHow come you always have to install some version of pytorch or tensor flow to run these ml models? When I'm only doing inference shouldn't there be easier ways of doing that, with automatic hardware selection etc. Why aren't models distributed in a standard format like onnx, and inference on different platforms solved once per platform?
- m3at 4y agoThat's done in professional contexts, when you only care about inference onnxruntime does the job well (including for coreml [1]). I imagine that here apple wants to highlight a more research/interactive use, for example to allow fine tuning SD on a few samples from a particular domain (a popular customization). [1] https://onnxruntime.ai/docs/execution-providers/CoreML-ExecutionProvider.html https://onnxruntime.ai/docs/execution-providers/CoreML-Execu...
- zitterbewegung 4y agoApple has their own mlmodel format but they can’t distribute this model as a direct download due to the models EULA. The first task is to translate the model.
- EMIRELADERO 4y agoWhat part of the SD license prohibits that?
- ronsor 4y agoNo part of it.
- judge2020 4y agoI mean, it is a legal time bomb in general[0], with a non-standard license that has special stipulations in an amendment. Do you really incur the weeks of lead time that it would take Legal to review the legality of redistributing this model? 0: https://github.com/CompVis/stable-diffusion/blob/main/LICENSE https://github.com/CompVis/stable-diffusion/blob/main/LICENS...
- darkteflon 4y agoFor the uninitiated, which MacOS GUI app is this library most likely to show up in first/best? DiffusionBee?
- pksebben 4y agoautomatic111's webui typically gets the most frequent updates. Middling easy to install.
- darkteflon 4y agoGreat, thank you. Look like there’s already a GH issue: https://github.com/AUTOMATIC1111/stable-diffusion-webui/issues/5309 https://github.com/AUTOMATIC1111/stable-diffusion-webui/issu...
- syspec 4y agoThere's also https://draw.nnc.ai/ https://draw.nnc.ai/ - which is an iOS / iPad app running Stable Diffusion. The author has a detailed blogpost outlining how he modified the model to use Metal on iOS devices. https://liuliu.me/eyes/stretch-iphone-to-its-limit-a-2gib-model-that-can-draw-everything-in-your-pocket/ https://liuliu.me/eyes/stretch-iphone-to-its-limit-a-2gib-mo...
- antal 4y agoYeah, that's what immediately came to mind for me as well. I don't know how similar/different the two solutions are, but it made me smile a bit that what Apple is showing off here has been already done by a single independent developer :)
- tamersalama 4y agoI can't get fine-tune the model ron Apple Silicon due to PyTorch supportability issues. I don't have high-hopes it will be supported. https://github.com/pytorch/pytorch/issues/77794 https://github.com/pytorch/pytorch/issues/77794 https://github.com/pytorch/pytorch/issues/77764 https://github.com/pytorch/pytorch/issues/77764
- tomr75 4y agoanyone know how to link this to a GUI?
- Viluskaran 4y ago8 gb ram
- Synaesthesia 4y agoWhat about it?
- wellthisisgreat 4y agoMacbook Air M1 / 16GB RAM took 3.56 to generate an image, this is pretty wild
- zimpenfish 4y ago> 3.56 to generate an image 3.56 seconds?
- wellthisisgreat 4y agoah 3.56 minutes, my mistake
- personjerry 4y agoCan't wait to see this integrated into automatic1111 so I can use it as a normie
- wilsongoode 4y agoI’ve been using InvokeAI: https://github.com/invoke-ai/InvokeAI https://github.com/invoke-ai/InvokeAI Great support for M1, basically since the beginning. The install is painless. Release video for InvokeAI 2.2: https://www.youtube.com/watch?v=hIYBfDtKaus https://www.youtube.com/watch?v=hIYBfDtKaus
- dustedcodes 4y agoWhat are some good resources to get into working with this and learning the basics around ML to get some fundamental understanding of how this works?
- videlov 4y agoI found the blog posts by Jay Alammar to be particularly good. Here are my starting suggestions (in this order) — https://jalammar.github.io/illustrated-word2vec/ https://jalammar.github.io/illustrated-word2vec/ https://jalammar.github.io/illustrated-transformer/ https://jalammar.github.io/illustrated-transformer/ https://jalammar.github.io/illustrated-bert/ https://jalammar.github.io/illustrated-bert/ https://jalammar.github.io/illustrated-stable-diffusion/ https://jalammar.github.io/illustrated-stable-diffusion/
- joss82 4y agoWould it be possible to run 2 SD instances in parallel on a single M1/M2 chip? One on the GPU and another on the ML core?
- cloogshicer 4y agoI think it's sad that Apple doesn't even give attribution to any of the authors. If you copy the Bibtex from this site, the Author field is just empty. Their names are also not mentioned anywhere on this site. This site is purely a marketing effort.
- rvz 4y ago> I think it's sad that Apple doesn't even give attribution to any of the authors. Pretty much like Stable Diffusion and the grifters using it in general and they will never credit the artists and images that they stole to generate these images.
- ClumsyPilot 4y agoDo your point is that Apple and those grifters are equally reputable? two wrongs don't make a right.
- rvz 4y agoI'm neither defending Apple or the grifters using Stable Diffusion in my comment. Both are as bad as each other, giving no attribution or credit.
- astrange 4y agoThis is sort of like if you learned English from reading a book and the author said they owned all your English sentences after that. Of course you can see the original images (https://rom1504.github.io/clip-retrieval/ https://rom1504.github.io/clip-retrieval/), it was legal to collect them (they used robots.txt for consent just like Google Image Search) and it was legal to do this with them (but not using US legal principles since it's made in Germany). "Crediting the artist" isn't a legal principle - it's more like some kind of social media standard which is enforced by random amateur artists yelling at you if you don't do it. It's both impossible (there are no original artists for a given output) and wouldn't do anything to help the main social issue (future artists having their jobs taken by AIs).
- noduerme 4y agoCan anyone explain in relatively lay terms how Apple's neural cores differ from a GPU? If they can run stable diffusion so much faster, which normally runs on a GPU, why aren't they used to run shaders for AAA games?
- Synaesthesia 4y agoThey're designed to run ML specific functions like matrix multiply and stuff. Nvidia has a similar idea in "tensor cores". I think because they're low but operations like 8 or 16 bit which is faster but too low res for GPU work.
- siraben 4y agoWhile running locally on an M1 Pro is nice, recently I've switched over to a Runpod[0] instance running Stable Diffusion instead. The main reasons being high workloads placed on the laptop degrade the battery faster and it takes ~40s to render a single image. On an A5000 it takes mere seconds to do 40 steps. The cost is around $0.2/hr. [0] https://runpod.io https://runpod.io