3 ms·
It's true that the programming model allows that, but underneath all threads within the warp will execute the same instructions. However if there is a branch, s
by gpuhacker 8y ago
It's true that the programming model allows that, but underneath all threads within the warp will execute the same instructions. However if there is a branch, some threads can be predicated so their instructions have no effect. This is called warp divergence and can become a performance issue. If possible branch only on threadidx using multiples of the warp size. There's a cool slide deck on implementing a parallel sum algorithm that explains this really well.
- dialecticDolt 8y agoDo you have any idea what the source of the slide deck could have been? It sounds very interesting and I'd love to see it.
- gpuhacker 8y agoSure! I meant this one: https://developer.download.nvidia.com/assets/cuda/files/reduction.pdf https://developer.download.nvidia.com/assets/cuda/files/redu...