3 ms·
Can you cite a reference explaining this ability to re-form new warps from existing threads? I ask because I’ve seen posts from NVIDIA support saying that dive
by subharmonicon 3y ago
Can you cite a reference explaining this ability to re-form new warps from existing threads?
I ask because I’ve seen posts from NVIDIA support saying that divergence is still very expensive and I’ve also seen benchmarks that force divergence in each warp by evenly splitting the warp, and the benchmarks result in 2x runtime when that happens vs. when the control-flow is dynamically uniform.
One thing to keep in mind is that even if you were to dynamically reform-warps, there’s still a potential expense because you then lose the advantage of doing things like accessing adjacent elements of data in adjacent threads. You’re bound to now have more bank conflicts, fewer memory accesses being coalesced, etc. Perhaps they do actually do this warp re-formation, but that itself does have this additional cost.
- winwang 3y ago"Divergence is still very expensive" is quite compatible with "Divergence is less expensive than before". Here's evidence (not proof) that Nvidia would remove the hard limit of warp divergence (or perhaps more "precisely", *a warp is always synchronous across its 32 threads with divergence"): https://developer.nvidia.com/blog/cooperative-groups/ https://developer.nvidia.com/blog/cooperative-groups/ I don't think it's misleading to talk about a "CUDA core" as a warp-wide processor, although it seems that Nvidia doubles the number (at least for gaming GPUs), presumably because of having both FP and INT pathways.
- subharmonicon 3y agoTheir “CUDA core” is not warp-wide, it’s a single lane. If you’re talking about FP32 rates, they double it because of FMA (floating-point multiply-accumulate). Everyone does that.
- deleted 3y ago[deleted]
- freeone3000 3y agothe SIMT section there is pretty telling — you can do it, if you explicitly account for it, and are willing to leave other threads in the dust potentially forever. It’s not quite the same thing as JMP and only seems to account for the data dependency case and not the if/else non-stream case. But I stand corrected — Volta and up have multiple PCs per warp.