8 ms·
Stable Diffusion on AMD RDNA3
- AstixAndBelix 4y agoDoes anyone know what's the current state of AMD's tools to migrate from CUDA? There's so much untapped potential with these cards, it's crazy that basically only gamers can make use of their competitive prices
- epmaybe 4y agoI don’t think there’s truly a competitor but opencl is the alternative to shoot for. Otherwise for machine learning purposes amd helps develop ROCm.
- pjmlp 4y agoOpenCL is hardly an alternative, plain old C, using compilation from source at runtime, with very basic tooling available. Versus a polyglot compiler infrastructure, IDE tooling that includes shader debugging, and a rich ecosytem of GPU based libraries. Even with SYSCL and SPIR-V, that has hardly improved, and while Intel bases oneAPI on top of SYSCL, that naturally also goes beyond the standard.
- amelius 4y agoShouldn't we have an API that can speak to both CUDA and opencl? Or is opencl sufficiently capable?
- pjmlp 4y agoNo it isn't because it lacks the polyglot infrastructure from CUDA, it has now SPIR-V but hardly anyone targets it as PTX gets used.
- snvzz 4y agoI understand AMD HiP is a CUDA clone, where library functions have the same syntax but with hip replacing cuda in the function names. Behind, it can use AMD and NVIDIA hardware alike. Thus, the idea is that through typically negligible effort porting to HiP, your code becomes vendor-independent. In practice, I do not know how true this is.
- my123 4y ago> Thus, the idea is that through typically negligible effort porting to HiP, your code becomes vendor-independent. Here, the big AMD mistake was to rename those function prefixes in the first place. It's a mistake that they could have avoided... What a lot of SW codebases did to support AMD (see PyTorch code notably): codebase is still CUDA, have the conversion pass to HIP done at build time. See https://github.com/ROCm-Developer-Tools/HIPIFY/blob/amd-staging/bin/hipify-perl https://github.com/ROCm-Developer-Tools/HIPIFY/blob/amd-stag... for the Perl script to do it. Then comes the problem of AMD not supporting ROCm HIP on most of their hardware or user base. On Windows, the ROCm HIP SDK is private and only available under NDA. This means that while you can use Blender w/ HIP on Windows, the Blender builds that you compile yourself will not be able to use ROCm HIP. On Linux, the supported GPUs are few and far between, Vega20 onwards are supported today. APUs, RDNA1, and lower end RDNA2 w/o unsupported hacks (6700 XT and below) are excluded.
- tormeh 4y agoIt's quite baffling. AMD is behaving like an incumbent trying to segment users etc. when they really should behave more like an upstart trying to make things easy. But their drivers for Linux are the best, so I don't think I'll switch to Nvidia...
- paulmd 4y ago> What a lot of SW codebases did to support AMD (see PyTorch code notably): codebase is still CUDA, have the conversion pass to HIP done at build time. This is sort of echoed in AMD's stance on FSR2/upscaling, where they have explicitly stated they will not support any API that allows plugging, regardless of whether the API is open-source or not, or who owns it, because a pluggable API might allow plugging proprietary implementations. Their opinion is their solution is what's best for everyone, in every situation, so you don't really need DLLs or pluggability because why would you want to plug something worse? FSR2 is the best for everyone and you should really just be compiling it directly into your application. https://youtu.be/8ve5dDQ6TQE?t=974 https://youtu.be/8ve5dDQ6TQE?t=974 (this of course also makes it impossible to update versions of FSR2 if better ones come out subsequently - you can't do the DLSS thing where you swap in newer DLLs (at your own risk of course) and benefit from later improvements to a modular grouping of code. You know, sort of the whole concept of libraries in the first place...) The HIP stuff is the same thing... AMD really wants you to convert once to HIP and be locked in forever, because Theirs Is The Best, Why Would You Need Anything Else? But of course HIP doesn't have a PTX-like concept so you really need to distribute as source and compile everything at runtime... because who would want library code or dynamic linking? Anyway, like, I know it's not really a shocker but the "we love open-source!" thing is a bit of an act. They love it when it's an angle for them as the underdog to leverage their way into marketshare... and as the underdog when it's not favorable for them (like FSR) they'll abandon their pro-freeness stance. And they too have their closed, proprietary technologies (like their CXL alternative that only works with their CPU+peripherals and nobody else can use) that they don't open up either. Nor is AMD racing to open up chipsets (like the NForce or Abit days) either, that's all locked down and proprietary too. I know that's not really a shocker when you put it like that, but, AMD really gets a ton of the benefit-of-the-doubt all the time. They have on multiple occasions shipped defective/marginal silicon at launch for example, and it all just gets brushed over and people forget all about it. Both Zen2 (low-quality silicon in the launch batch meant chips were missing advertised boost clocks by 10%+) and RDNA1 had massive incidents of the community downplaying very real problems because AMD Is Good Now, many of the affected users never had their problems resolved and they just kinda sighed and lived with it or sold the hardware and bought something better, and the fans swept it all under the rug and never talked of it again. Same for pandemic profiteering (while Intel cut prices), etc. There's just a ton of shit that people bend over backwards to find justifications for with AMD that just wouldn't fly with more reputable vendors.
- schmorptron 4y agoDo you have an opinion on the new openCL implementation that recently got merged into mesa? It doesn't touch on tooling or the other points you mentioned, but performance seems to be pretty good! https://www.phoronix.com/news/Rusticl-2022-XDC-State https://www.phoronix.com/news/Rusticl-2022-XDC-State
- ColonelPhantom 4y agoWhat do you mean by a polyglot compiler infrastructure? Are you referring to the fact that CUDA source is single-file (your host and device code are in the same compilation unit?) Or do you mean that you can ship the same binary to different GPU architectures? SYCL solves the first issue, and SPIR-V solves the second one. (OpenCL mostly avoids the issue in general though by making you ship source which is then compiled by the driver, but SPIR-V allows you to ship a 'binary' instead). No clue as for debugging and IDE tooling, but I did find a rocgdb binary on my Linux ROCm installation (which is for HIP, not SYCL). No clue what oneAPI offers for debugging. Furthermore, Clang (and hence clangd) speaks HIP and I think SYCL too. So the non-runtime IDE tooling should work. Finally, a lot of GPU libraries are I think available for ROCm/HIP too. It's unfortunate that the HIP stack sucks enormously in other ways.
- marcyb5st 4y agoLast time I seriously checked (6 months ago or so) ROCm was still a far cry from CUDA. Set up was a mess, support was hit and miss, some operations were not particularly performante compared to the CUDA counterparts. Additionally, there are Tensorflow and probably PyTorch forks that should work with it, but they lag behind the official repositories quite a bit. I hope that now that generative AI is becoming mainstream AMD steps up their game both on their consumer and professional lineups. If I were to buy a video card right now ( mostly for gaming+ML hobbies projects + running stable diffusion) I wouldn't pick AMD because I could do just 1/3 of my use cases properly without headaches (gaming).
- rowanG077 4y agoOpenCL works pretty well. Can't say I notice large gaps of performance between CUDA and openCL for my hpc work.
- lalaland1125 4y agoHave you done any benchmarks with vulkan?
- rowanG077 4y agoNo I haven't used vulkan for compute.
- my123 4y agoThankfully for a good chunk of number crunching that works fine. But the other side of the coin is notably AI workloads. There's no OpenCL or Vulkan standard for exposing matrix units, only vendor specific ones. For OpenCL: cl_qcom_ml_ops (Qualcomm) notably, for Vulkan: VK_NV_cooperative_matrix (NVIDIA)
- CodeArtisan 4y agoThe performances on Blender3d are atrocious, the RX 7900 XTX is noticeably slower than a RTX 3060.
- snvzz 4y agoLatest Blender release does not have the optimization work in yet. AIUI, what's in current git master is very different.
- ColonelPhantom 4y agoA big part of the reason is that Blender on Nvidia supports hardware accelerated ray tracing using OptiX. HIP-RT exists, but is not used in Blender yet. I think the Intel oneAPI backend for Arc GPUs also misses RT acceleration. AMD claims to have HIP-RT working internally, but not yet suitable for posting publically. Intel is planning it, I think. Both should land around Blender 3.6, if I'm not mistaken. If you take the raw FLOPS, CUDA (not OptiX) and HIP are actually nearly equivalent in performance last I remember. I think RDNA2 just does "more with less", at least in terms of gaming performance per FLOP (e.g. due to the huge cache).
- ggerganov 4y ago> There has also been a wide variety of accuracy-degrading performance optimizations like Xformers and Flash Attention, which are great tools if you are open to trading accuracy for performance .. I wasn't aware that Flash Attention trades accuracy for performance. Either I have a wrong understanding of what FA is, or this statement is not fully accurate. Either way - looks like great work
- marcyb5st 4y agoFrom the flash attention paper: We also extend FlashAttention to block-sparse attention, yielding an approximate attention algorithm that is faster than any existing approximate attention method. So I assume they are using the approximate version as they also have an exact version.
- ggerganov 4y agoThanks for that - I have missed the block-sparse extension of the algorithm when I first read about it. And indeed this seems to be what the author means.
- imhoguy 4y agoAny chance to get SD running on mobile Ryzen APU e.g. Ryzen Pro 4750U (Renoir)?
- delijati 4y agoShort answer no. Long answer "in theory" yes. I tried this [1] but gave up as building rocm + deps takes up to 6h :/ Official statement [2] [1] https://github.com/xuhuisheng/rocm-build https://github.com/xuhuisheng/rocm-build [2] https://github.com/RadeonOpenCompute/ROCm/issues/1587 https://github.com/RadeonOpenCompute/ROCm/issues/1587
- nicolaslem 4y agoFor anyone on Arch, there is a third-party repository called arch4edu[0] that provides up to date builds of ROCm and its dependencies. On my iGPU, OpenCL sometimes works, sometimes crashes. Even finding a list of supported hardware is close to impossible. The whole situation is just ridiculous and makes AMD look bad. [0] https://github.com/arch4edu/arch4edu https://github.com/arch4edu/arch4edu
- my123 4y agoAMD doesn't actually care. For them, GPGPU is a pro level feature not worth supporting on most customer GPUs. They are doing much more feature segmentation than NVIDIA ever did.
- sosborn 4y agoThis worked for me with a 5700xt: https://www.travelneil.com/stable-diffusion-windows-amd.html https://www.travelneil.com/stable-diffusion-windows-amd.html
- negativegate 4y agonod-ai/SHARK from the original submission is by far the fastest way I've found to run Stable Diffusion on a 5700 XT. For 50 iterations: * ONNX on Windows was 4-5 minutes * ROCm on Arch Linux was ~2.5 minutes * SHARK on Windows is ~30 seconds
- lalaland1125 4y agoI really wish more GPU libraries had focused on vulkan instead of CUDA ...
- rowanG077 4y agoIt's one of the reasons Nvidia is basically untouchable at this time. The AI field willingly enslaved itself to NVidia.
- my123 4y agoIt's because NVIDIA actually cared and AMD does not where it matters (customer HW). Openness is totally secondary to functional. You know, the same kind of reasons as of why Linux on the desktop is not a mass market thing compared to Windows for a very long time.
- rowanG077 4y agoOpenCL is totally functional, on AMD and even iGPU Intel. The reason Nvidia won was because they made it easier. And the AI people ate it up. The tooling Nvidia offers is second to none. But you can build almost anything CUDA does with openCL. It's simply harder to do. The AI crowd cared more about that then the impact of tieing the entire ecosystem to a single company. Who knows what openCL might have been if it would be the premier implementation language. I'd wager it would have gotten a LOT more love.
- deleted 4y ago[deleted]
- CamperBob2 4y agoWhen I catch myself writing a sentence like It's simply harder to do in order to promote or justify an alternative engineering approach, I try to, well, catch myself before making a really weak argument in favor of an inferior solution. Making life easy for developers is important.
- chem83 4y ago> SHARK is an open source cross platform (Windows, macOS and Linux) Machine Learning Distribution packaged with torch-mlir (for seamless PyTorch integration), LLVM/MLIR for re-targetable compiler technologies along with IREE (for efficient codegen, compilation and runtime) and Nod.ai’s tuning. IREE is part of the OpenXLA Project Google has been doing a good job advancing the IREE ML compiler project, which I think is what will bring other hw platforms like AMD and Intel to the ML game. Industry only has to benefit from increased hardware portability.
- tehsauce 4y ago“There has also been a wide variety of accuracy-degrading performance optimizations like Xformers and Flash Attention, which are great tools if you are open to trading accuracy for performance” This is incorrect. Those optimizations do identical computations, but leverage memory bandwidth on the gpu more effectively. So there is no accuracy tradeoff there.
- magic_at_nodai 4y agoHere are a list of potential issues https://github.com/AUTOMATIC1111/stable-diffusion-webui/discussions/2592#discussioncomment-4024360 https://github.com/AUTOMATIC1111/stable-diffusion-webui/disc... That said we (Nod.ai team) will add support for xformers soon so you can opt in for xformers anyway.
- thunkshift1 4y agoCan someone explain what exactly does nod.ai do? Its not clear at all from their page
- brokenmachine 4y agoCan anyone point me to some examples of what I, as a techie, might want to actually use AI for? Some simple hobby projects?