4 ms·
Isn't this a bit dated? I certainly agree that the shared memory semantics are a critical distinction, but in addition to the 256Kb per SM shared memory, NVIDIA
by sephamorr 5y ago
Isn't this a bit dated? I certainly agree that the shared memory semantics are a critical distinction, but in addition to the 256Kb per SM shared memory, NVIDIA Volta GPUs have 128kB L1$ per SM, and a unified 6MB L2$. I don't think these caches are entirely /not worth discussing/.
- dragontamer 5y ago> I don't think these caches are entirely /not worth discussing/ Hmmm... NVidia's is clearly aiming at just punching through the memory-bandwidth problem with GDDR6x (2-bits per clock tick since its got 4-level encoding). That's the thing, NVidia isn't really pushing memory bandwidth or size limits on the L2 or even L1 cache IMO. Even the 128kB L1$ per SM is only roughly the size of the SM's register space. Their most interesting move really is GDDR6x, which is the brute-force way to solve that problem. -------- AMD's "Infinity Cache" on RDNA2 is 128MB of L3$, but AMD is using only standard GDDR6 (1-bit per clock tick transferred). AMD's RDNA2 is very strange: L0, L1, L2, and L3 caches, when AMD GCN was just L1 and L2 layers of cache. That "infinity cache" is worth talking about I guess... its large enough to be relevant in a number of gaming situations. ------ I guess AMD and NVidia are both using HBM at the high end for 1TBps to 2TBps bandwidths. But those chips aren't in the consumer realm anymore. The ultimate brute force solution: spend more money. You're right in that the L1 and L2 caches (and L0 and L3 caches of AMD) probably do affect performance in real ways.
- shaklee3 5y agoRegisters and cache are used together, though. And shared memory is the same as the L1 cache, but manually controlled.
- dragontamer 5y agoNVidia __shared__ takes from L1 cache. AMD does not. L1 and __shared__ are different pools on GCN, CDNA, and RDNA architectures. I believe shared is actually higher bandwidth than L1 on AMD systems, especially with atomics. > Registers and cache are used together, though But not for the same purpose, or the same way as CPUs. The cache is non-cohesive, large amounts of the cache are "K$", constant space that's non-cohesive. Etc. etc. Its a bit different. Some level of caching will improve effective memory bandwidth, but it seems like the GPU's primary strategy is to "float" register loads. GPUs are still an in-order processor but... the load-register assembly instructions clearly execute in an async-like manner. That load-register assembly instruction could be from L1 cache, it could be from GDDR6x, it could be from another GPU over NVLink / NVSwitch, or it could be even from PCIe (!!), being stored on the CPU's DDR4 RAM all the way across the motherboard. Doesn't matter: the load-register instruction will appropriately start loading the data, and will indicate to the core when the register is ready (and the GPU core will task-switch to other kernels while waiting for that request to be completed).