4 ms·
ROCm just doesn't have the support these days. Supposedly people are working on getting stable diffusion working on it. https://www.videogames.ai/2022/11/06/Sta
by stuckinhell 3y ago
ROCm just doesn't have the support these days.
Supposedly people are working on getting stable diffusion working on it.
https://www.videogames.ai/2022/11/06/Stable-Diffusion-AMD-GPU-ROCm-Linux.html https://www.videogames.ai/2022/11/06/Stable-Diffusion-AMD-GP...
But it's just too much of investment for me on something that MAY work.
I ended up just buying a 4080rtx
- pja 3y agoStable Diffusion worked fine for me with Rocm & my RX580 (once I had compiled a custom torch library IIRC). But I don’t know whether it works with the more recent RDNA2 cards.
- pohuing 3y agoIt works without a hitch on my rx6900xt. Only pain was getting the amdpro drivers
- thewataccount 3y agoYeah that's kinda what I mean. I've heard to not even consider AMD for ML applications specifically ROCM, so I'm curious if these chips will use ROCM as their primary api or not. It "technically" works, but their own examples crash, you get a fraction of the performance you'd expect for the level of hardware you have, chicken&egg problem with little other software having good support of it, etc.
- delusional 3y agoI'm running stable diffusion on my 7900xtx and it's working fine. I had to screw around a little bit to get the newest ROCm and torch libraries since they aren't packed on my OS, but it wasn't that bad. I made a docker image if anybody is struggling to get it working: https://hub.docker.com/r/delusional/sd-rx7900xtx https://hub.docker.com/r/delusional/sd-rx7900xtx
- thewataccount 3y agoOut of curiosity, how many it/s do you get with DPM2 at 512x512 with a batch size of 1, and then the it/s for whatever the max batch size you can fit?
- delusional 3y agoIt does 15-16 it/s with euler a and 2.16 it/s at 8 batchsize (the max in automatic1111), and that's only using 15GiB of vram.
- brucethemoose2 3y agoStable Diffusion works fine on rocm (and intel OpenVINO), the issue is out-of-the-box support in popular UIs. TBH the whole space is kinda a mess. Tons of optimizations (like most ML compilers), even on Nvidia cards, are left on the table because the SD UI devs just dont have the throughout or motivation to implement them. At the other end, hardware makers, ml compiler devs, researchers and such are making quick demos, but are not making any integration attempts for popular frameworks. There is no one in the middle, so we are stuck with PyTorch eager mode and a perception that it only works on big Nvidia GPUs.