5 ms·
The Debian ROCm Team [1] has made quite a bit of progress in getting the ROCm stack into the official Archive. Most components are already packaged, the next b
by ckastner 3y ago
The Debian ROCm Team [1] has made quite a bit of progress in getting the ROCm stack into the official Archive.
Most components are already packaged, the next big target is adding support to the PyTorch package.
Many of the packages are older versions; this is because getting broad coverage was prioritized. The other next big target that is currently being worked on is getting full ROCm 5.7 support.
I fully expect Debian 13 (trixie) to come with full ROCm support out-of-the-box, and as a consequence, also derivatives to have support (Ubuntu above all). In fact, there will almost certainly be backports of ROCm 5.7 to Debian 12 (bookworm) within the next few months, so one will be able to just
$ sudo apt-get install pytorch-rocm
One current obstacle is infrastructure: the Debian build and CI infrastructures (both hardware and software) were not designed with GPUs in mind. This is also being worked on.
Edit: forgot to say that the CI infra that the Team is setting up here tests all of these packages on consumer cards, too. So while there may not be official support for most of these, upstream tests passing on the cards within the infra should be a good indication for practical support.
[1] https://salsa.debian.org/rocm-team/ https://salsa.debian.org/rocm-team/
- avcxz 3y agoI'd also like to point out that ROCm has been packaged for Arch Linux since the beginning of 2023, with efforts starting since March 2020 [1]. Currently on Arch Linux you can run the following successfully: $ sudo pacman -S python-pytorch-rocm Arch Linux even has ROCm support with blender. [1] https://github.com/rocm-arch https://github.com/rocm-arch
- alright2565 3y agoHope you don't mind, but I have a rant I need to get out. I decided to give this another try now that you've mentioned it. Let's get things started the way the arch wiki suggests: $ sudo pacman -S rocm-hip-sdk $ /opt/rocm/bin/clinfo ERROR: clGetPlatformIDs(-1001) $ sudo /opt/rocm/bin/clinfo ... Board name: AMD Radeon RX 6600 XT ... Ok, I wonder what's wrong. maybe it's this? https://stackoverflow.com/questions/4959621/error-1001-in-clgetplatformids-call https://stackoverflow.com/questions/4959621/error-1001-in-cl... Nope. Anything about this on the arch wiki? Nope This bug report[2] from 2021? Maybe I need to update my groups. [2]: https://github.com/RadeonOpenCompute/ROCm/issues/1411 https://github.com/RadeonOpenCompute/ROCm/issues/1411 $ ls -la /dev/kfd crw-rw-rw- 1 root render 237, 0 Sep 26 20:33 /dev/kfd $ sudo usermod -aG render $(whoami) $ # relogin $ /opt/rocm/bin/clinfo ERROR: clGetPlatformIDs(-1001) Ok, I'm a pretty advanced linux user, I'll just jump right in: $ strace /opt/rocm/bin/clinfo ... openat(AT_FDCWD, "rusticl.icd", O_RDONLY|O_NONBLOCK|O_CLOEXEC|O_DIRECTORY) = -1 ENOENT (No such file or directory) Apparently I have some leftover environment variables (OCL_ICD_VENDORS) from last time I spent half a day trying to get this to work. I can fix that. After all, it'd be entirely unreasonable to expect rocm to give me a better error, like "Could not open opencl icd `rusticl.icd`". Success: $ /opt/rocm/bin/clinfo Number of platforms: 1 ... Board name: AMD Radeon RX 6600 XT Well, let's run some apps! $ darktable -d opencl ... [dt_opencl_device_init] DEVICE: 0: 'gfx1032' PLATFORM NAME & VENDOR: AMD Accelerated Parallel Processing, Advanced Micro Devices, Inc. ... PHI node has multiple entries for the same basic block with different incoming values! %967 = phi float [ %largephi.extractslice0, %sw.default ], [ %largephi.extractslice055, %sw.bb667 ], [ %largephi.extractslice059, %sw.bb663 ], [ %largephi.extractslice063, %sw.bb659 ], [ %largephi.extractslice067, %sw.bb655 ], [ %largephi.extractslice071, %sw.bb646 ], [ %largephi.extractslice075, %_Z4fmodff.exit16 ], [ %largephi.extractslice079, %_Z4fmodff.exit13 ], [ %largephi.extractslice083, %_Z4fmodff.exit ], [ %largephi.extractslice087, %sw.bb562 ], [ %largephi.extractslice091, %sw.bb555 ], [ %largephi.extractslice095, %sw.bb533 ], [ %largephi.extractslice099, %if.then502 ], [ %largephi.extractslice0103, %if.else517 ], [ %largephi.extractslice0107, %if.then456 ], [ %largephi.extractslice0111, %if.else471 ], [ %largephi.extractslice0115, %if.then393 ], [ %largephi.extractslice0119, %if.else408 ], [ %largephi.extractslice0123, %if.then338 ], [ %largephi.extractslice0127, %if.else353 ], [ %largephi.extractslice0131, %if.then283 ], [ %largephi.extractslice0135, %if.else298 ], [ %largephi.extractslice0139, %if.then224 ], [ %largephi.extractslice0143, %if.else241 ], [ %largephi.extractslice0147, %sw.bb193 ], [ %largephi.extractslice0151, %sw.bb180 ], [ %largephi.extractslice0155, %sw.bb168 ], [ %largephi.extractslice0159, %sw.bb158 ], [ %largephi.extractslice0163, %sw.bb147 ], [ %largephi.extractslice0167, %if.then116 ], [ %largephi.extractslice0171, %if.else131 ], [ %largephi.extractslice0175, %sw.bb71 ], [ %largephi.extractslice0179, %sw.bb ], [ %largephi.extractslice0183, %if.end ], [ %largephi.extractslice0187, %if.end ], [ %largephi.extractslice0191, %if.end ], [ %largephi.extractslice0195, %if.end ], [ %largephi.extractslice0199, %if.end ] label %if.end %largephi.extractslice0183 = extractelement <4 x float> %div, i64 0 %largephi.extractslice0191 = extractelement <4 x float> %div, i64 0 in function blendop_Lab LLVM ERROR: Broken function found, compilation aborted! [1] 27586 IOT instruction (core dumped) darktable -d opencl uh that's great. Maybe blender? It worked! Not too bad for 2 minutes render: https://i.imgur.com/FD1SsQG.png https://i.imgur.com/FD1SsQG.png What about pytorch? It prompted this whole thing anyway: $ sudo pacman -S python-pytorch-rocm python-torchvision $ python neural_style/neural_style.py eval --content-image ../../2min.png --model ./saved_models/mosaic.pth --output-image out.png --cuda 1 [1] 32471 segmentation fault (core dumped) python neural_style/neural_style.py eval --content-image ../../2min.png $ sudo dmesg --follow [ 2467.536713] python[33309]: segfault at 68 ip 00007f12c5504d5d sp 00007ffc8f539c20 error 4 in libamdhip64.so.5.6.31062[7f12c541e000+357000] likely on CPU 14 (core 7, socket 0) [ 2467.536727] Code: ec 78 48 89 bd 78 ff ff ff 64 48 8b 04 25 28 00 00 00 48 89 45 c8 31 c0 85 f6 0f 88 09 03 00 00 48 8b 85 78 ff ff ff 48 63 de <48> 8b 50 68 48 8b 40 70 48 89 85 70 ff ff ff 48 29 d0 48 c1 f8 03 uh oh. Maybe I can crack some passwords? $ hashcat -m 0 -a 0 -o cracked.txt target_hashes.txt /usr/share/dict/american-english ... hiprtcCompileProgram(): HIPRTC_ERROR_COMPILATION error: unknown argument: '-flegacy-pass-manager' 1 error generated when compiling for gfx1032. * Device #1: Kernel /usr/share/hashcat/OpenCL/shared.cl build failed. Well, so much for that. Best I can get to work with rocm is 1/4 apps.
- slavik81 3y agoThis is why Christian and I have invested so much effort into the CI system for Debian. There needs to be a clear accounting of what works and what doesn't for every library on every architecture.
- slavik81 3y agoIt's too late to edit, but I should add that the RX 6600 XT is not officially supported by the upstream ROCm project. It's not clear to me that the experience would be better on any other distro. That's where having public test logs would be valuable.
- blihp 3y agoIIRC, nothing below the 6800 is supported by ROCm... so the lion's share of their installed base of the 6000 series is excluded from 'official' support. nVidia's compute drivers support all of their devices and have across multiple generations, AMD's support only the low volume devices and drop support for older generations seemingly almost as fast as they are released.
- imtringued 3y agoI could get hashcat to work with poor performance but then the computer was unusable.
- lhl 3y agoOne of your problems might be that gfx1032 is not supported by AMD's ROCm packages, which has a laughably short list of supported hardware: https://rocm.docs.amd.com/en/latest/release/gpu_os_support.html#linux-supported-gpus https://rocm.docs.amd.com/en/latest/release/gpu_os_support.h... The normal workaround is to assign the closest architecture, eg gfx1030, so `HSA_OVERRIDE_GFX_VERSION=10.3.0` might help Also, it looks like some of your tested projects are OpenCL? For me, I do something like: `yay -S rocm-hip-sdk rocm-ml-sdk rocm-opencl-sdk` to cover all the bases. My recent interest has been LLMs and this is my general step by step for those (llama.cpp, exllama) for those interested: https://llm-tracker.info/books/howto-guides/page/amd-gpus https://llm-tracker.info/books/howto-guides/page/amd-gpus I didn't port the docs back in, but also here's a step-by-step w/ my adventures getting TVM/MLC working w/ an APU: https://github.com/mlc-ai/mlc-llm/issues/787 https://github.com/mlc-ai/mlc-llm/issues/787 From my experience, ROCm is improving, but there's a good reason that Nvidia has 90% market share even at big price premiums. EDIT: apparently Darktable and Blender have OpenCL issues that are fixed in the just released 5.7: https://github.com/ROCm-Developer-Tools/clr/issues/3 https://github.com/ROCm-Developer-Tools/clr/issues/3
- slavik81 3y ago> One current obstacle is infrastructure: the Debian build and CI infrastructures (both hardware and software) were not designed with GPUs in mind. This is also being worked on. To be more direct, one thing we lack is funding. AMD has provided RDNA 2 and RDNA 3 GPUs for the Debian CI, but to fill out the rest of the architecture matrix I have been personally buying GPUs. That's been sufficient for covering most architectures, but we will need a sponsor if we are to acquire CDNA 2 and CDNA 3 hardware. Our goal is to cover every modern discrete AMD GPU architecture on the CI. At the moment, that would be Navi 33, Navi 32, Navi 31, Navi 24, Navi 23, Navi 22, Navi 21, Navi 14, Navi 12, Navi 10, Aldebaran, Arcturus, Vega 20, Vega 10, and (maybe) Polaris. I have been very successful at bringing the AMD GPU libraries to architectures that are not officially supported upstream. Unfortunately, I can't afford to keep buying systems out of my personal funds. I have personally spent ~7k USD on hardware for the CI and I have been offered reimbursement from the Debian project for my next ~5k USD in spending. That has given us a good foundation, but we could do more to improve hardware support if we had more funding available. Please consider donating to the Debian Project [1] if you wish to support their efforts. [1]: https://www.debian.org/donations https://www.debian.org/donations
- ssivark 3y ago> but to fill out the rest of the architecture matrix I have been personally buying GPUs. That's been sufficient for covering most architectures, but we will need a sponsor if we are to acquire CDNA 2 and CDNA 3 hardware. This seems like the kind of thing that AMD should be providing (or at least sponsoring) as a matter of principle — regardless of whether it can be funded in other ways I.e. if anyone in AMD cared this problem would be solved trivially. The fact that you are funding it out of pocket is seriously calling into question AMD’s commitment. What am I missing here?
- neilv 3y agoI think Debian is one of the first places AMD should be looking to fund for ROCm. Especially if there's already motivated capable volunteer labor, and all they need is equipment, and maybe a devrel point of contact. The cost seems like a few peanuts dug out of the sofa cushions, on a strategic push like this.