4 ms·
NVIDIA is intentionally being obtuse and frankly dishonest calling what’s effectively a vector lane a “core” and similarly using “thread” in “SIMT” to mean the
by subharmonicon 3y ago
NVIDIA is intentionally being obtuse and frankly dishonest calling what’s effectively a vector lane a “core” and similarly using “thread” in “SIMT” to mean the execution of one of those vector lanes.
Yes, their architecture is different from many in that they support a separate program counter per lane (which is why they feel justified in calling this a “thread”), but ultimately it’s the rate and throughput of ALUs that matter.
- freeone3000 3y agoIt’s not even a seperate PC per lane, you only get that per block — lane level execution goes to an execution mask lut per insn. Branchy code that’s not branchy uniformly in the block executes a lot of noops.
- sakras 3y agoAs of Volta, they have independent PCs with a warp optimizer that dynamically groups threads with the same program counter, so branches aren’t nearly as bad as they used to be.
- subharmonicon 3y agoCan you cite a reference explaining this ability to re-form new warps from existing threads? I ask because I’ve seen posts from NVIDIA support saying that divergence is still very expensive and I’ve also seen benchmarks that force divergence in each warp by evenly splitting the warp, and the benchmarks result in 2x runtime when that happens vs. when the control-flow is dynamically uniform. One thing to keep in mind is that even if you were to dynamically reform-warps, there’s still a potential expense because you then lose the advantage of doing things like accessing adjacent elements of data in adjacent threads. You’re bound to now have more bank conflicts, fewer memory accesses being coalesced, etc. Perhaps they do actually do this warp re-formation, but that itself does have this additional cost.
- winwang 3y ago"Divergence is still very expensive" is quite compatible with "Divergence is less expensive than before". Here's evidence (not proof) that Nvidia would remove the hard limit of warp divergence (or perhaps more "precisely", *a warp is always synchronous across its 32 threads with divergence"): https://developer.nvidia.com/blog/cooperative-groups/ https://developer.nvidia.com/blog/cooperative-groups/ I don't think it's misleading to talk about a "CUDA core" as a warp-wide processor, although it seems that Nvidia doubles the number (at least for gaming GPUs), presumably because of having both FP and INT pathways.
- subharmonicon 3y agoTheir “CUDA core” is not warp-wide, it’s a single lane. If you’re talking about FP32 rates, they double it because of FMA (floating-point multiply-accumulate). Everyone does that.
- deleted 3y ago[deleted]
- freeone3000 3y agothe SIMT section there is pretty telling — you can do it, if you explicitly account for it, and are willing to leave other threads in the dust potentially forever. It’s not quite the same thing as JMP and only seems to account for the data dependency case and not the if/else non-stream case. But I stand corrected — Volta and up have multiple PCs per warp.