4 ms·
The 1024 threads of a warp/block/whatever or just the current threads or what?
by DSingularity 5y ago
The 1024 threads of a warp/block/whatever or just the current threads or what?
- dragontamer 5y ago> The 1024 threads of a warp/block/whatever In CUDA terminology: the block can access __shared__ memory together. Different blocks are locked out of seeing other block's __shared__ memory. A block can be up to 32-warps working together (aka 1024 threads), or it could be as small as 1-warp (aka 32-threads). A warp could be 1-thread in some cases. -- Finally, a grid in CUDA is synchronized across kernel calls. You are 100% certain that all threads of the grid are not running before the kernel_launch<<<x, y, z, stream>>(foo bar)... and you can be 100% certain that all threads of the grid are done after the cudaStreamSynchronize(stream) call. --------- Warps: Largely about very low-level details such as branch divergence. Warps take if/else/for loops together in practice, so you need to think about warps when you think about optimal utilization. In the past 5 years or so, warp-level programming has become more popular, but warp-level stuff is pretty rare and only should be reserved once other, easier, optimizations have taken place. Blocks: Coordinated across __shared__ memory. Because blocks are controlled by the programmer, its easier to think about than warps... but you need to be careful and put the right OpenCL barriers() or CUDA __syncthreads() in the right places. Grid: Coordinated across cudaStreamSynchronize / CUDA-streams. Ad-hoc coordination can be done with atomics to global memory, but this is very slow if memory barriers are involved. (Relaxed atomics are pretty fast though)
- DSingularity 5y agoThanks!