7 ms·
Zen, CUDA, and Tensor Cores, Part I: The Silicon
- downvotetruth 2y agoI refused to buy the so determined defective chips even if they represented better value because if the intent was truly to try and max yield then there should be for Ryzen for example good 7 core versions with only 1 core that was found to be defective. Since no 7 core zens exist, then at least some of the CPUs with 6 core CCDs have intentionally had 1 of the cores destroyed for reasons unknown, which could be to meet volume targets. If this is because for Ryzen the cores can only be disabled in pairs, then it boggles my mind that it would not be economic given the $ diff of tens to hundreds of dollars between the 6 and 8 core versions that is does not make sense to add the circuits to allow each core to be individually fused off and allow further product differentiation, especially considering how much effort and # of SKUs have been put forth with the frequency binning in AM4 (5700x, 5800, 5800x, 5800xt, etc.), rather than bigger market segmentation jumps.
- AnthonyMouse 2y ago> if the intent was truly to try and max yield then there should be for Ryzen for example good 7 core versions with only 1 core that was found to be defective. Since no 7 core zens exist There are Zen processors that use 7 cores per CCD, e.g. Epyc 7663, 7453, 9634. The difference between Ryzen and Epyc is the I/O die. The CCDs are the same so that's presumably where they go. Another reason you might not see this on the consumer chips is that they have higher base clocks. If you have a CCD where one core is bad and another isn't exactly bad but can't hit the same frequencies as the other six, it doesn't take a lot of difference before it makes more sense to turn off the slowest than lower the base clock for the whole processor. 6 x 4.7GHz is faster than 7 x 4.0GHz, much less 7 x 2.5GHz. In theory you could let that one core run at a significantly lower speed than the others, but there is a lot of naive software that will misbehave in that context. Whereas the base clock for the Epyc 9634 is 2.25GHz, because it has twelve 7-core CCDs so it's nearly 300W, and doesn't want to be nearly 1300W regardless of whether or not most of the cores could do >4GHz.
- downvotetruth 2y agoTo correct the example for the Epyc line, models appears to exist with 1 through 8 cores available except for 5.
- AnthonyMouse 2y agoThe Epyc models with lower core counts per CCD probably don't exist because of yields though. The 73F3 has two cores per CCD, so with eight CCDs it only has 16 cores. The 7303 also has 16 cores but two CCDs, so all eight cores per CCD are active. The 73F3 costs more than five times as much. That's weird if the 73F3 is the dumping ground for broken dice. Not so weird when you consider that it has four times as much L3 cache and higher clock speeds. The extra cores in the 73F3 aren't necessarily bad, they're disabled so the others can have their L3 cache and so they can pick the two cores from each CCD that hit the highest clock speeds. Doing that is expensive, especially if the other cores aren't all bad, but then you get better performance per core. Which some people will pay a premium for, so they offer models like that even if yields are good and there aren't that many CCDs with that many bad cores. At which point your premise is invalid because processors are being sold with cores disabled for performance reasons rather than yield reasons.
- downvotetruth 2y ago> they're disabled so the others can have their L3 cache and so they can pick the two cores from each CCD that hit the highest clock speeds what or where does that follow from? One can take a CCD with 2+ cores and pin a process to a set (of the fastest) cores based on profiling the cores and those 2+ cores could use the L3 cache as needed; disabling cores at the hardware level is the waste as if they were not disabled, then that would allow other processes to be able to benefit from more than 2 cores to run when desired. The latter point of disabling cores for "better [frequency] performance per core Which some people will pay a premium for" is dubious especially for the Epyc server line. If that were true, then there should at least be 4 core or fewer SKUs for desktop Ryzen variant where apps like games are more likely to benefit from the higher clock.
- tverbeure 2y agoThat’s first sentence is a spectacular non-sequitur.
- MobiusHorizons 2y agoI would guess that there is a desire to not create too many product tiers. I believe 6 core parts are made from 2 3-core CCXs, (rather than 4 and 2) so only one core is disabled per ccx.
- cinnamonteal 2y agoCurrent Ryzen and EPYC processors have 8 core CCXs. The 6 core parts used to be as you described, but are now a single CCX. The Zen C dies have two CCXs, but they are still 8 core CCXs, and are always symmetrical in core count. The big exception is that the new Zen 5 Strix Point chip has a 4 core CCX for the non-C cores. I think the Zen 4 based Z1 has a similar setup but don't remember and couldn't quickly find the actual information to confirm.
- wtallis 2y agoThe Ryzen Z1 was a weird one: two Zen4 cores plus four Zen4c cores all in one cluster, sharing the same 16MB L3 cache.
- Symmetry 2y agoIt would be sort of cool if they could do direct to consumer sales with every core going at whatever its maximum speed is or turned off if to disrupted. But that's not something you could do through existing distribution channels, everyone presumes a fairly limited number of SKUs.
- diabllicseagull 2y agoIt was a good read. I wonder what hot takes he'll have in the second part if any.
- fulafel 2y agoThe answer to the leading question "What’s the difference between a Zen core, a CUDA core, and a Tensor core?" is not covered in Part 1, so you may want to wait if this interests you more than chip layouts.
- raphlinus 2y agoHere's my quick take. A top of the line Zen core is a powerful CPU with wide SIMD (AVX-512 is 16 lanes of 32 bit quantities), significant superscalar parallelism (capable of issuing approximately 4 SIMD operations per clock), and a high clock rate (over 5GHz). There isn't a lot of confusion about what constitutes a "core," though multithreading can inflate the "thread" count. See [1] for a detailed analysis of the Zen 5 line. A single Granite Ridge core has peak 32 bit multiply-add performance of about 730 GFLOPS. Nvidia, by contrast, uses the marketing term "core" to refer to a single SIMD lane. Their GPUs are organized as 32 SIMD lanes grouped into each "warp," and 4 warps grouped into a Streaming Multiprocessor (SM). CPU and GPU architectures can't be directly compared, but just going by peak floating point performance, the most comparable granularity to a CPU core is the SM. A warp is in some ways more powerful than a CPU core (generally wider SIMD, larger register file, more local SRAM, better latency hiding) but in other ways less (much less superscalar parallelism, lower clock, around 2.5GHz). A 4090 has 128 SMs, which is a lot and goes a long way to explaining why a GPU has so much throughput. A 1080, by contrast, has 20 SMs - still a goodly number but not mind-meltingly bigger than a high end CPU. See the Nvidia Ada whitepaper [2] for an extremely detailed breakdown of 4090 specs (among other things). A single Nvidia 4090 "core" has peak 32 bit multiply-add performance of about 5 GFLOPS, while an SM has 640 GFLOPS. I don't know anybody who counts tensor cores by core count, as the capacity of a "core" varies pretty widely by generation. It's almost certainly best just to compare TFLOPS - also a bit of a slippery concept, as that depends on the precision and also whether the application can make use of the sparsity feature. I'll also note that not all GPU vendors follow Nvidia's lead in counting individual SIMD lanes as "cores." Apple Silicon, by contrast, uses "core" to refer to a grouping of 128 SIMD lanes, similar to an Nvidia SM. A top of the line M2 Ultra contains 76 such cores, for 9728 SIMD lanes. I found Philip Turner's Metal benchmarks [3] useful for understanding the quantitative similarities and differences between Apple, AMD, and Nvidia GPUs. [1]: http://www.numberworld.org/blogs/2024_8_7_zen5_avx512_teardown/ http://www.numberworld.org/blogs/2024_8_7_zen5_avx512_teardo... [2]: https://images.nvidia.com/aem-dam/Solutions/Data-Center/l4/nvidia-ada-gpu-architecture-whitepaper-V2.02.pdf https://images.nvidia.com/aem-dam/Solutions/Data-Center/l4/n... [3]: https://github.com/philipturner/metal-benchmarks https://github.com/philipturner/metal-benchmarks
- kvemkon 2y ago> Each of the tiles on the CPU side is actually a Zen 4 core, complete with its dedicated L2 cache. Perhaps, it could be more interesting to compare without L2 cache.
- Symmetry 2y agoOr maybe a CUDA core versus one of Zen's SIMD ports.
- adrian_b 2y agoThe L2 really belongs to the core, a comparison without it does not make much sense. The GPU cores (in the classic sense, i.e. not what NVIDIA names as "cores") also include cache memories and also local memories that are directly addressable. The only confusion is caused by the fact that first NVIDIA, and then ATI/AMD too, have started to use an obfuscated terminology where they have replaced a large number of terms that had been used for decades in the computing literature with other terms. For maximum confusion, many terms that previously had clear meanings, like "thread" or "core", have been reused with new meanings and ATI/AMD has invented a set of terms corresponding to those used by NVIDIA but with completely different word choices. I hate the employees of NVIDIA and ATI/AMD who thought that it is a good idea to replace all the traditional terms without having any reason for this. The traditional meaning of a thread is that for each thread there exists a distinct program counter a.k.a. instruction pointer, which is used to fetch and execute instructions from a program stored in the memory. The traditional meaning of a core is that it is a block that is equivalent with a traditional independent processor, i.e. equivalent with a complete computer minus the main memory and the peripherals. A core may have only one program counter, when it can execute a single thread at a time, or it may have multiple program counters (with associated register sets) when it can execute multiple threads, using either FGMT (fine-grained multithreading) or SMT (simultaneous multithreading). The traditional terms were very clear and they have direct correspondents in GPUs, but NVIDIA and AMD use other words for those instead of "thread" and "core" and they reuse the words "thread" and "core" for very different things, for maximum obfuscation. For instance, NVIDIA uses "warp" instead of "thread", while AMD uses "wavefront" instead of "thread". NVIDIA uses "thread" to designate what was traditionally named the body of a "parallel for" a.k.a. "parallel do" program structure (which when executed on a GPU or multi-core CPU is unrolled and distributed over cores, threads and SIMD lanes).
- Darulquran-123 2y ago[flagged]
- paulmd 2y agoyou can calculate the area of the tensor and raytracing units by measuring+comparing die sizes between the nearest 20-series and 16-series chips. Contrary to the assumptions a lot of people made from the cartoon diagrams, it's actually relatively small, together they make up approximately 18% of the cluster area and it's below 10% of the chip as a whole. The area is roughly 2/3rds tensor unit area and 1/3 raytracing unit area, so RT is around 3% of total chip area and tensor is around 6%. https://old.reddit.com/r/hardware/comments/baajes/rtx_adds_195mm2_per_tpc_tensors_125_rt_07/ https://old.reddit.com/r/hardware/comments/baajes/rtx_adds_1... This could have changed somewhat in newer releases, but probably not too drastically, since NVIDIA has never really increased raw ray performance since the 20-series launch. And while there have been a few raytracing features around the edges, raster and cache have been bumped significantly too (notably, ampere got dual-issue fp32 pipelines... which didn't really work out for NVIDIA that well either!) so honestly there's a reasonable chance it's slightly less in subsequent architectures.