3 ms·
4090 has 16384 cuda cores, so there's that!
by Keyframe 2y ago
4090 has 16384 cuda cores, so there's that!
- atq2119 2y agoNvidia's marketing is misleading. Those "cuda cores" are more SIMD lanes than cores. Number of SMs is a more appropriate equivalent to CPU core count.
- Keyframe 2y agoyou're right
- dahart 2y agoAre you sure that isn’t what @mullingitover meant? > Number of SMs is a more appropriate equivalent to CPU core count. What do you mean by this? Why should an SM be considered equivalent to a CPU core? An SM can do 128 simultaneous adds and/or multiplies in a single cycle, where a CPU core can do, what, 2 or maybe 4? Obviously depends on the CPU / core / hyperthreading / # math pipelines / etc., but the SM to CPU-core ratio of the number of simultaneous calculation is in the double digits. It’s a tradeoff where the GPU has some restrictions in return for being able to do many multiples more at the same time. If you consider an SM and a CPU equivalent, then the SM’s perf can exceed the CPU core by ~2 orders of magnitude — is that the comparison you want? If you consider a GPU thread lane and a CPU thread lane equivalent, the the GPU thread lane is slower and more restricted. Neither comparison is apples to apples, CPUs and GPUs are made for different workloads, but arguing that an SM is equivalent to a CPU core seems equally or more “misleading” when you’re leaving out the tradeoff. I’d argue that comparing SMs to cores is misleading and that it makes more sense to compare chips is by their thread counts. Or, don’t compare cores at all and just look at the performance in, say, FLOPS.
- atq2119 2y agoA single Zen5 core can do 32 single precision FMAs per clock. That's using SIMD, but so is Nvidia for all intents and purposes. Those "cuda cores" aren't truly independent: when their execution diverges, masking is used pretty much like you'd do in CPU SIMD. A lot of the control logic is per-SM or perhaps per-SIMD unit -- there are multiple of those per SM. You could perhaps make a case that it's the individual SIMDs which correspond to CPU cores (that makes the flops line up even more closely). It depends on what the goal of the comparison is.
- deleted 2y ago[deleted]
- Dylan16807 2y ago> An SM can do 128 simultaneous adds and/or multiplies in a single cycle, where a CPU core can do, what, 2 or maybe 4? https://images.nvidia.com/aem-dam/Solutions/geforce/news/rtx-40-series-vram-video-memory-explained/nvidia-ada-lovelace-gpu-architecture-streaming-multiprocessor.png https://images.nvidia.com/aem-dam/Solutions/geforce/news/rtx... An SM is split into four identical blocks, and I would say each block is roughly equivalent to a CPU core. It has a scheduler, registers, 32 ALUs or FPUs, and some other stuff. A CPU core with two AVX-512 units can do several integer operations plus 32 single-precision operations (including FMA) per cycle. Not 2 or 4. An older CPU with 2-3 AVX2 units could fall slightly behind, but it's pretty close. That doesn't factor in the tensor units, but they're less general purpose, and CPUs usually put such things outside the cores. I would say an SM is roughly equivalent to four CPU cores.
- dahart 2y agoYeah I totally forgot to consider CPU SIMD. Brain fart. The other comment corrected me too, you’re both right. When I said 2 or 4, I was thinking of the SISD math pipe and not AVX instructions. Yes, considering CPU SIMD, maybe comparing a CPU core to a CUDA warp makes some sense in some situations. The peak FLOPS rate is still so much higher on Nvidia though, that the comparison hardly makes sense. So yeah like I and the other commenter mentioned, it depends entirely on what comparison is being made.