4 ms·
winwang’s comment is correct, yours is wrong. “cuda core” refers to one lane within the SIMT/SIMD ALUs. These lanes within a SMSP don’t execute independently.
by rnrn 2y ago
winwang’s comment is correct, yours is wrong.
“cuda core” refers to one lane within the SIMT/SIMD ALUs. These lanes within a SMSP don’t execute independently.
The term SMSP is definitely used for nvidia’s architecture :
https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-decoder https://docs.nvidia.com/nsight-compute/ProfilingGuide/index....
> smsp
> Each SM is partitioned into four processing blocks, called SM sub partitions. The SM sub partitions are the primary processing elements on the SM.
(Note that this kernel profiling guide doesn’t use the term “cuda cores” at all)
Also there are 128 “cuda cores” per SM in 4090, not 64 :
https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvidia-ada-gpu-architecture.pdf https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvid...
> Each SM in AD10x GPUs contain 128 CUDA cores