4 ms·
Oooh, I didn’t know PTX was an intermediate representation and explicitly documented as such, I really thought it was the actual assembly ran by the chips…
by navaati 6y ago
Oooh, I didn’t know PTX was an intermediate representation and explicitly documented as such, I really thought it was the actual assembly ran by the chips…
- my123 6y agoYou can get the GPU-targeted assembly (sometimes called SASS by NVIDIA) through specifically compiling to a given GPU then using nvdisasm, which also has a very terse definition of the underlying instruction set in the docs (https://docs.nvidia.com/cuda/cuda-binary-utilities/index.html#instruction-set-ref https://docs.nvidia.com/cuda/cuda-binary-utilities/index.htm...). But it's one way only, NVIDIA ships a disassembler, but explicitly doesn't ship an assembler.
- dogma1138 6y agoPTX is a virtual ISA and the reason why ROCm is doomed to fail well beyond its horrendous bugs. ROCm produces hardware specific binaries which not only means you need to produce binaries for multiple GPUs but you also don’t have any guarantee for forward compatibility. A CUDA binary from 10 years ago will still run today on modern hardware, ROCm breaks compatibility between minor releases sometimes and it’s often not documented.