5 ms·
I’m not using Frontier, but I am using Setonix which is a large AMD cluster being rolled out in Australia. All of AMD’s teaching materials are about ROCm so thi
by torrance 4y ago
I’m not using Frontier, but I am using Setonix which is a large AMD cluster being rolled out in Australia. All of AMD’s teaching materials are about ROCm so this is very much how they’re expecting it to be used.
The real pain for us is that there’s no decent consumer grade chips with ROCm compatibility for us to do development on. AMD have made it very clear they only care about the data centre hardware when it comes to ROCm, but I have no idea what kind of developer workflow they’re expecting there.
- uniqueuid 4y agoInteresting. So what is your workflow right now?
- torrance 4y agoDevelop against CUDA locally. Port my kernels to ROCm, and occupy a whole HPC node for debugging and performance tuning for a week. It’s terrible. Edit: I should say that their recommendation is to write the kernels in ‘hip’ which is supposed to be their cross device wrapper for both cuda or ROCm. I’m writing in Julia however so that’s not possible.
- claforte 4y agoThe AMD software stack has been behind for a long time but I feel like we're finally catching up. I heard that HIP (and hopefully the rest of ROCM) is now supported on the RX6800XT consumer GPU... maybe that could help? BTW my team at AMD has been using Julia for ML workloads for a while. We should get in touch - maybe some of the lessons we learn can be useful to you. My email is claforte. The domain I'm sure you can guess. ;-)
- vchuravy 4y agoIf you are using Julia I would recommend looking at AMDGPU.jl and (pluging my own project here) KernelAbstractions.jl
- claforte 4y agoBTW have you tried `KernelAbstractions.jl`? With it you can write code once that will run reasonably fast on AMD or NVIDIA GPUs or even on CPU. One of our engineers just started using it and is pleased with it - apparently the performance is nearly equivalent to native CUDA.jl or AMDGPU.jl, and the code is simpler.
- sorenjan 4y agoCan you write SYCL code and compile it to ROCm for production?
- tkinom 4y agoNV21? https://www.phoronix.com/scan.php?page=news_item&px=Radeon-ROCm-5.0 https://www.phoronix.com/scan.php?page=news_item&px=Radeon-R...
- JonChesterfield 4y agoThe rocm stack will run on non-datacentre hardware in YMMV fashion. A lot of the llvm rocm development is done on consumer hardware, the rocm stack just isn't officially tested on gaming cards during the release cycle. In my experience codegen is usually fine and the Linux driver a bit version sensitive.
- tormeh 4y agoIt's bare pickings, but there are chips: https://docs.amd.com/bundle/Hardware_and_Software_Reference_Guide/page/Hardware_and_Software_Support.html https://docs.amd.com/bundle/Hardware_and_Software_Reference_...
- dragontamer 4y agoVega64 or Vega56 seems to work pretty well with ROCm in my experience. Hopefully AMD gets the Rx 6800xt working with ROCm consistently, but even then, the 6800xt is RDNA2, while the supercomputer Mx250x is closer to the Vega64 in more ways. So all in all, you probably want a Vega64, Radeon VII, or maybe an older MI50 for development purposes.
- slavik81 4y ago> Hopefully AMD gets the Rx 6800xt working with ROCm consistently I am a maintainer for rocSOLVER (the ROCm LAPACK implementation) and I personally own an RX 6800 XT. It is very similar to the officially supported W6800. Are there any specific issues you're concerned about? I know the software and I have the hardware. I'd be happy to help track down any issues.
- dragontamer 4y agoThat's good to hear. I might be operating off of old news. But IIRC, the 6800 wasn't well supported when it first came out, and AMD constantly has been applying patches to get it up-to-speed. I wasn't sure what the state of the 6800 was (I don't own it myself), so I might be operating under old news. As I said a bit earlier, I use the Vega64 with no issues (for 256-thread workgroups. I do think there's some obscure bug for 1024-thread workgroups, but I haven't really been able to track it down. And sticking with 256-threads is better for my performance anyway, so I never really bothered trying to figure this one out)
- slavik81 4y agoNavi 21 launched in November 2020 but it only got official support with ROCm 5.0 in February 2022. With respect to your issue running 1024 threads per block, if you're running out of VGPRs, you may want to try explicitly specify the max threads per block as 1024 and see if that helps. I recall that at one point the compiler was defaulting to 256 despite the default being documented as 1024.
- 4y ago
- eslaught 4y agoI'm surprised you're not using HIP? At least in my experience it seems like HIP is the go-to system for programming the AMD GPUs, in large part because of CUDA compatibility. You can mostly get things to work with a one-line header change [1]. (I work for a DOE lab but views are my own, etc.) [1] As an example, see the approach in: https://github.com/flatironinstitute/cufinufft/pull/116 https://github.com/flatironinstitute/cufinufft/pull/116
- pmarcelll 4y agoHIP is just the programming language/runtime, ROCm is the whole software stack/platform.