4 ms·
When people say CUDA is Nvidia specific, I don't think they realize _how_ specific. It's not just a matter of having the API implementation. CUDA does things l
by jms55 3y ago
When people say CUDA is Nvidia specific, I don't think they realize _how_ specific. It's not just a matter of having the API implementation.
CUDA does things like depend on specific scheduler behavior for GPU threads in order to guarantee forward progress, allowing more efficient single-pass computation routines. Or allowing CUDA kernels to launch additional kernels or allocate GPU memory from within the kernel, without host communication. Or checking the GPU architecture and performing specific micro-optimizations designed around that hardware's internal design.
GPU's are nothing like CPU's, in that the x86 architecture is fairly stable. There's not a _ton_ of difference between different CPUs. Sure, performance characteristics and cache sizes might be different, but generally they have the same instruction set. Each GPU generation is basically a completely different architecture, much less between GPU companies. That's (partly) why Vulkan is so complicated - it tries to support the lowest spec 2013 mobile GPU, and the highest spec 2023 desktop GPU in one single API.
- nologic01 3y agoThis is the result of GPU computationally intensive tasks not being as simply defined as the CPU workloads. It was always thus for the various accelerators and HPC designs. They must make design choices that favor a relatively narrow range of computational patterns whereas the CPU will happily chew through a wider range of instruction sequences with the same performance. The corresponding GPU software complexity is something that feels not quite appreciated. Yet if we are going to see the massive adoption of GPU's that the market projects, with multiple providers of hardware, we'll need some serious advances of the software side.
- pjmlp 3y agoI think we need more stuff like Futhark or StarLisp. Being able to describe GPU computations in a declarative way, with a driver similar to an SQL query optimizer, distributing the load.