4 ms·
One thing I don’t understand is how the architecture of Apple Silicon is different from NVidia’s. Looking at this quote: > the Nvidia H100 GPU has 132 SMs wit
by guidedlight 3y ago
One thing I don’t understand is how the architecture of Apple Silicon is different from NVidia’s.
Looking at this quote:
> the Nvidia H100 GPU has 132 SMs with 64 cores per SM, totalling a whopping 8448 cores.
8448 cores sure sounds impressive. But the Apple M2 Ultra only has 76 cores?!
How can the NVidia H100 GPU have over 110x more cores? Clearly it doesn’t have 110x more performance over the M2 Ultra, so what is going on here?
- kevingadd 3y agoNVIDIA's SMs are most comparable to the 'CUs' on AMD GPUs or cores on Apple GPUs, generally speaking. The "cores" are subsets of the SM that perform individual operations, IIRC. See this diagram from an nvidia blog post: https://developer-blogs.nvidia.com/wp-content/uploads/2021/guc/raD52-V3yZtQ3WzOE0Cvzvt8icgGHKXPpN2PS_5MMyZLJrVxgMtLN4r2S2kp5jYI9zrA2e0Y8vAfpZia669pbIog2U9ZKdJmQ8oSBjof6gc4IrhmorT2Rr-YopMlOf1aoU3tbn5Q.png https://developer-blogs.nvidia.com/wp-content/uploads/2021/g... ( https://developer.nvidia.com/blog/nvidia-ampere-architecture-in-depth/ https://developer.nvidia.com/blog/nvidia-ampere-architecture... )
- subharmonicon 3y agoNVIDIA is intentionally being obtuse and frankly dishonest calling what’s effectively a vector lane a “core” and similarly using “thread” in “SIMT” to mean the execution of one of those vector lanes. Yes, their architecture is different from many in that they support a separate program counter per lane (which is why they feel justified in calling this a “thread”), but ultimately it’s the rate and throughput of ALUs that matter.
- freeone3000 3y agoIt’s not even a seperate PC per lane, you only get that per block — lane level execution goes to an execution mask lut per insn. Branchy code that’s not branchy uniformly in the block executes a lot of noops.
- sakras 3y agoAs of Volta, they have independent PCs with a warp optimizer that dynamically groups threads with the same program counter, so branches aren’t nearly as bad as they used to be.
- subharmonicon 3y agoCan you cite a reference explaining this ability to re-form new warps from existing threads? I ask because I’ve seen posts from NVIDIA support saying that divergence is still very expensive and I’ve also seen benchmarks that force divergence in each warp by evenly splitting the warp, and the benchmarks result in 2x runtime when that happens vs. when the control-flow is dynamically uniform. One thing to keep in mind is that even if you were to dynamically reform-warps, there’s still a potential expense because you then lose the advantage of doing things like accessing adjacent elements of data in adjacent threads. You’re bound to now have more bank conflicts, fewer memory accesses being coalesced, etc. Perhaps they do actually do this warp re-formation, but that itself does have this additional cost.
- winwang 3y ago"Divergence is still very expensive" is quite compatible with "Divergence is less expensive than before". Here's evidence (not proof) that Nvidia would remove the hard limit of warp divergence (or perhaps more "precisely", *a warp is always synchronous across its 32 threads with divergence"): https://developer.nvidia.com/blog/cooperative-groups/ https://developer.nvidia.com/blog/cooperative-groups/ I don't think it's misleading to talk about a "CUDA core" as a warp-wide processor, although it seems that Nvidia doubles the number (at least for gaming GPUs), presumably because of having both FP and INT pathways.
- subharmonicon 3y agoTheir “CUDA core” is not warp-wide, it’s a single lane. If you’re talking about FP32 rates, they double it because of FMA (floating-point multiply-accumulate). Everyone does that.
- mr_toad 3y agoFor one thing you can use the H100 to heat a room - it uses more than 10x the power of an M2 Ultra.
- winwang 3y agoConsider AMD Epyc 7742 vs an A100 -- 225W vs 400W (according to some TDP numbers). That's 1.8x-2.0x increase in energy, but something like 10x-100x more integer operations per second (depending on size of integers, depending on if you count the tensor cores or not).