5 ms·
Area-wise I'm pretty sure Apple's P-cores are in line with Intel and AMD - including L2 cache and other shared logic, M1's P-core was 4mm^2 per core, M2's P-cor
by brigade 4y ago
Area-wise I'm pretty sure Apple's P-cores are in line with Intel and AMD - including L2 cache and other shared logic, M1's P-core was 4mm^2 per core, M2's P-core is about 5mm^2 [1], Zen 3 and 4 were about 4mm^2 [2], and Intel is 7mm^2 for Golden Cove [3]
[1] https://www.semianalysis.com/p/apple-m2-die-shot-and-architecture https://www.semianalysis.com/p/apple-m2-die-shot-and-archite...
[2] https://www.hwcooling.net/en/zen-4-architecture-chip-parameters-and-ipc-of-amds-new-core/ https://www.hwcooling.net/en/zen-4-architecture-chip-paramet...
[3] https://twitter.com/locuza_/status/1453524285260247046 https://twitter.com/locuza_/status/1453524285260247046
- dragontamer 4y agoM2 P-core is 5mm^2 on 3nm process. Zen3 is 4mm^2 on 7nm process... a process that uses 4x the area per transistor than the 3nm process.
- brigade 4y agoM1, M2, and Zen 4 are all on TSMC 5nm TSMC only just started mass production of 3nm like last week, nothing has shipped yet.
- dragontamer 4y agoMost of the benchmarks I'm familiar with are M1/M2 vs Zen3. I don't think people have comprehensively tested Zen4 yet. Zen3 was 7nm for sure. So AMD Zen3 ~4mm^2 on 7nm is ~364 million transistors, while Apple's ~5mm^2 on 5nm is ~856 million transistors (https://en.wikichip.org/wiki/File:5nm_densities.svg https://en.wikichip.org/wiki/File:5nm_densities.svg). With M1/M2 in the ~5mm^2 area or so, I definitely argue that Zen3 cores are 1/2 the size of M1/M2 cores, transistor-for-transistor. Its a fat core. Maybe it will work thanks to how advanced processes are getting. Maybe this will encourage others (ie: Intel) to experiment with larger cores as well. Its hard for me to say, but I do welcome the benchmarks.
- brigade 4y agoWell yeah, area budgets per-core have never shrunk linearly with transistor density; it serves a wider range of use cases to balance beefing up cores with adding more of them. Like, Intel 7 is 20x denser than Intel 32nm, but a Sandy Bridge core is less than 3x larger than Golden Cove. Also the L2 cache and shared logic make up a larger percentage of M1/M2 per-core at >45%; that's only 20% of the per-core area for Zen 3... if you include LLC that doubles the per-core Zen3 area but only adds 30% for M2... Point is that M1/M2 and Zen 4 show that the per-core area budget within the same process is now similar across Apple and AMD, not an order of magnitude different. It used to be an order of magnitude different, like back on 32nm Apple A6 was about 8 mm^2/core and Sandy Bridge was 18.5 mm^2/core, or 30 mm^2 including LLC that A6 didn't have.
- dragontamer 4y ago> Also the L2 cache and shared logic make up a larger percentage of M1/M2 per-core at >45%; AMD's L2 cache is 1MB on Zen4, and L3 cache is like 4MB. (and L2 cache compares to Apple L1 cache, while L3 AMD cache is LLC / Comparable to Apple's L2 cache). AMD's L2 cache is per-core. AMD's L3 cache is decentralized last-level, is 32MB for 8-cores (equivalent to Apple's 16MB for 4 cores). Except... AMD's cores have 2 threads on them while Apple only has 1. I think my overall point is clear: Apple's cores are abnormally large. AMD / Intel have smaller cores (and larger caches). This is _despite_ shoving 2-thread per core on AMD/Intel through SMT or Hyperthreading. Remember that only one thread gets that HUGE core on Apple. Its very, very unusual. Even POWER10 (which has oversized cores) allows 8x SMT (8 threads per core) to compensate for its oversized nature.
- brigade 4y agoI mean, I was trying to be fair by counting L2 and associated logic. If you only count down to L1 (and no, Zen4's 15-cycle latency L2 is not comparable to Apple's L1 that achieves a 3-4 cycle latency; Zen4's combined L2+L3 averages close to Apple's 18-cycle L2), then a Zen4 core only takes 72% of that 3.84 mm^2, or 2.76 mm^2. M2's P-core is estimated at 2.756 mm^2 if scaled to match M1's density, or 2.519 mm^2 if you accept Apple marketing's scaling. And the M1's P-core was 2.281 mm^2. Hyperthreading barely costs any area, but anyway I guess you can say that thanks to that plus the clock speed advantage, Zen4 gets like 15% more performance per mm^2 than M2 P-cores? That's not a massive improvement by any measurement. (as an aside: if annotations of Zen4 I've seen are correct, its branch predictor has almost as much SRAM as the entire µop+L1i+L1d caches. Which... actually I can completely believe of TAGE)