3 ms·
How much of this is because CUDA is designed for GPU execution and because the GPU ISA isn’t a stable interface? E.g. new GPU instructions can be utilized by ne
by variadix 2y ago
How much of this is because CUDA is designed for GPU execution and because the GPU ISA isn’t a stable interface? E.g. new GPU instructions can be utilized by new CUDA compilers for new hardware because the code wasn’t written to a specific ISA? Also, don’t people fine tune GPU kernels per architecture manually (either by hand or via automated optimizers that test combinations in the configuration space)?
- dragontamer 2y agoNVidia PTX is a very stable interface. And the PTX to SASS compiler DOES a degree of automatic fine tuning between architectures. Nothing amazing or anything, but it's a minor speed boost that has made PTX just a easier 'assembly-like language' to build on top of.
- janwas 2y agoMy understanding is that there is a lot of hand-writing (not just fine-tuning) going on. AFAIK CuDNN and TensorRT are written directly as SASS, not CUDA. And the presence of FP8 in H100, but not A100, would likely require a complete rewrite.
- dragontamer 2y agoCub, thrust and many other libraries that make those kernels possible don't need to be rewritten. When you write a merge sort in CUDA, you can keep it across all versions. Maybe the new instructions can improve a few corner cases, but it's not like AVX to AVX512 where you need to rewrite everything. Ex: https://github.com/NVIDIA/cub/blob/main/cub/device/device_merge_sort.cuh https://github.com/NVIDIA/cub/blob/main/cub/device/device_me...
- janwas 2y agoI agree not everything needs to be rewritten. And neither does code using an abstraction such as Highway, so we can stop beating that dead horse.