3 ms·
Sometimes you write some heavily tuned code in a high level language like C++ that you know could be translated into very specific GPU assembly, then find that
by david-gpu 1y ago
Sometimes you write some heavily tuned code in a high level language like C++ that you know could be translated into very specific GPU assembly, then find that the compiler isn't producing the exact assembly that you had in mind.
When you talk to the computer team about it they may offer a range of solutions, some of which may not be applicable to open source code. Picture proprietary #pragmas, intrinsics, or whatnot. What do you do? You can't ship a high performance library that doesn't deliver high performance. It is then when you rely on things like function names to enable specific code transformations that can't be used in general because they would sometimes break third party code.
I never worked on Cutlass, but this is the sort of thing that is done in the real world.
There is nothing nefarious about this sort of optimizatkon. People comparing this to cheating on benchmarks by rendering lower quality images are not on the right track.
- School-Cotton 1y agoWhy wouldn’t you just use inline assembly in that case?
- krapht 1y agoThere is no way to write inline SASS (assembly equivalent) for CUDA code. You can inline PTX, but PTX is a high level bytecode designed to be portable.
- dahart 1y agoPTX is sometimes referred to as assembly, and it is an ISA, much lower level than C++. When people talk about writing inline assembly for CUDA, they mean PTX, and the C++ compiler’s inline assembly “asm” statement assumes PTX. For the most part you have much more control and ability to produce exactly the SASS you want when using PTX compared to when using C/C++. https://developer.nvidia.com/blog/understanding-ptx-the-assembly-language-of-cuda-gpu-computing/ https://developer.nvidia.com/blog/understanding-ptx-the-asse... https://eunomia.dev/others/cuda-tutorial/02-ptx-assembly/ https://eunomia.dev/others/cuda-tutorial/02-ptx-assembly/