3 ms·
Most GPU's only have so much compute cores. Even if you have a million threads, you might only have a few hundreds cores to execute them. If you have heavy bra
by maeln 3y ago
Most GPU's only have so much compute cores. Even if you have a million threads, you might only have a few hundreds cores to execute them.
If you have heavy branching you might slow down the whole lane (which size can vary from 8 to 64 execution most of the time). Which at this scale still do make a big difference. Although for small branching, masking help avoid too much slowdown.
- WJW 3y agoTo expand on this, even the latest Nvidia 4090 card "only" has about 16k cores. That's a lot, but it's hardly "millions" either.
- ben-schaaf 3y agoTo expand further the 16k cores aren't real cores. What people would normally consider a core is what Nvidia calls a "streaming multiprocessor" (SM), of which the 4090 has 128. These cores each have 128 "cuda cores", analogous to a SIMD lane - although not quite since "cuda cores" are themselves SIMD. To simplify: Each 128 SMs can run a different program and each of their 128 "cuda cores" can do ~single-cycle (small) matrix operations.
- 6keZbCECT2uB 3y agoI'm not sure about the 4090, but most of the GPUs I use have a warp size of 32, and warp divergence affects only up to those 32 threads. If you have a branch and all threads agree, you only walk down one branch. My mental model is a bit more like you have collections of warps in a block, and all warps in a block get scheduled onto an SM. Different GPU architectures allow for different numbers of warps to be simultaneously active or inactive, and each warp has its own instruction pointer and can be suspended while waiting for things like memory. I found the picture on pg 22 here really helpful: https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia-ampere-architecture-whitepaper.pdf https://images.nvidia.com/aem-dam/en-zz/Solutions/data-cente... Note that although there's 4 schedulers, on the A100, they don't dispatch every cycle iirc.
- ben-schaaf 3y agoThose are tensor cores, not cuda cores. They're used for AI rather than general compute/shaders. The 4090 has 512 of those. Correct me if I'm wrong, but as far as I can tell tensor cores are just accelerators. They can't do general compute: no branch or jump.
- 6keZbCECT2uB 3y agoThe tensor core accelerates mostly matrix operations and is the big block you can see has 4 per SM. Cuda core refers to the thread per SM, which you can see as FP32 or INT32 units, so there are (32*4) per SM on that diagram. Like you said, tensor core is similar to a special purpose ALU and is at a lower level of abstraction than something with an instruction pointer.