3 ms·
Why is it so hard to develop a CUDA-like to AMD-opcode compiler? If it wasn't taking the tech community so long I would imagine it would not be harder than por
by mFixman 3y ago
Why is it so hard to develop a CUDA-like to AMD-opcode compiler?
If it wasn't taking the tech community so long I would imagine it would not be harder than porting GCC to a new architecture.
- dotnet00 3y agoI think the issues are mostly relating to getting the optimizer right. After all, not much of a point to GPU acceleration if it isn't meaningfully faster than CPU. A lot of these compatibility layers have this issue, DirectML, ZLUDA, etc. GPUs tend to expect the compiler to bake in things like instruction reordering/out of order execution for optimal performance. The other challenge is that there isn't an "AMD-opcode", each generation tends to change around the opcodes a bit, so you want to compile to an intermediate representation which the driver would ingest and compile down to what the GPU uses. NVIDIA uses PTX for this, it works very well. AMD's ROCm doesn't use an IR, it compiles the code for each architecture, which means they have a very limited support window for consumer GPUs (to limit binary size). OpenCL and Vulkan supports SPIR-V as an IR, but IIRC, the OpenCL SPIR-V on AMD is very buggy, and Vulkan SPIR-V is very different. This has other side effects, like BLAS library support range. Hard for an open source community to justify putting in tons of effort into optimizing an entire BLAS library for each generation, when it'll only be supported for ~4 years (so, by the time you're finishing up with driver stability and library optimization, you're probably already at least a quarter of the way through the support period).
- mFixman 3y agoNeat, that makes sense. Still, giant missed opportunity for AMD not to focus on this when NVIDIA has an almost monopoly on massive GPGPUs.