4 ms·
I mean, I was trying to be fair by counting L2 and associated logic. If you only count down to L1 (and no, Zen4's 15-cycle latency L2 is not comparable to Apple
by brigade 4y ago
I mean, I was trying to be fair by counting L2 and associated logic. If you only count down to L1 (and no, Zen4's 15-cycle latency L2 is not comparable to Apple's L1 that achieves a 3-4 cycle latency; Zen4's combined L2+L3 averages close to Apple's 18-cycle L2), then a Zen4 core only takes 72% of that 3.84 mm^2, or 2.76 mm^2. M2's P-core is estimated at 2.756 mm^2 if scaled to match M1's density, or 2.519 mm^2 if you accept Apple marketing's scaling.
And the M1's P-core was 2.281 mm^2.
Hyperthreading barely costs any area, but anyway I guess you can say that thanks to that plus the clock speed advantage, Zen4 gets like 15% more performance per mm^2 than M2 P-cores? That's not a massive improvement by any measurement.
(as an aside: if annotations of Zen4 I've seen are correct, its branch predictor has almost as much SRAM as the entire µop+L1i+L1d caches. Which... actually I can completely believe of TAGE)
- dragontamer 4y agoL2 on Zen3/Zen4 is __PER CORE__. That's private memory, inside of each core, for operations. If you cut out L2 cache, the Zen3 / Zen4 core shrinks significantly. As per the Zen4 article you had: > The L2 cache in the cores has increased from 512 kB to 1MB, which also increases the occupied area a bit, but the cores are still smaller overall than Zen 3 on 7nm thanks to the 5nm process. The area including L2 cache is 3.84mm² The 3.84mm^2 figure _INCLUDES_ 1MB of L2 cache. If you wanna cut that out, doing so will damage your own argument, as the Zen4 core will shrink rather dramatically. (Especially with that "unshrinkable SRAM" argument you're trying to make). ----------- Look, I don't even know where you're going with this. It shouldn't be a surprise to anybody that a 8x wide M2 core with like 800-entry reorder buffer and 600-entry register file will be bigger than a 6x wide AMD Zen4 core with like 400-entry reorder buffer and like 300-entry register file. M2 was designed to be big, fat, and wide in execution. That's just how it works. And its a very interesting (arguably brilliant) tradeoff. But if you look at the damn chip, its just bigger. That's what happens when you add more stuff to a core, the core gets larger. AMD on the other hand, is narrower (especially on a per-thread basis: 2-threads fit on this smaller core), and instead spends way more transistors on L2 cache. Maybe _YOU_ don't like the tradeoff (Sure, I agree that AMD's L2 cache is 15 cycles latency), but maybe throughput is more important and you're overly focused on unimportant / hypothetical latency issues (the entire L2 cache can be accessed at full throughput IIRC). At the end of the day, we gotta get the devices and then benchmark them with real programs to really see what the sum of all these tradeoffs are. But I don't think there's much argument to be had here that the M1/M2 Apple cores are just bigger. I mean... we know the buffer sizes. We all know Apple's buffers are just bigger. ------- https://pbs.twimg.com/media/FbXiXZeaAAA2reZ?format=jpg&name=large https://pbs.twimg.com/media/FbXiXZeaAAA2reZ?format=jpg&name=... Look at this. As I stated before, the biggest "penalty" to the AMD Zen4 core is the uop cache (which is unnecessary in the Apple chip). You can just... look at the damn die shot. If you want to argue about legitimate space-saving ability of ARM systems, focus on _THAT_ part of the chip. You're talking about all sorts of things that aren't actually helping your side of the argument.
- brigade 4y ago> But I don't think there's much argument to be had here that the M1/M2 Apple cores are just bigger > You pretty much can make 2 cores fit inside of the M1 core My entire point has simply been debunking this. I'm pointing out that Apple, Intel, and AMD have similarly large area budgets for their big cores. Like, looking at actual chips produced on the same TSMC processes, you cannot fit two Zen4 cores inside of the space taken by one M1 or M2 P-core. You cannot fit two Zen2 or Zen3 cores within the space taken by one A12Z P-core. All of them have a somewhat similar per-core area budget, with difference in L2/LLC cache tradeoffs being the biggest differentiator in area. And yes, even outside of cache they make different tradeoffs with what they spend the area on. Zen4 spends area on 512b registers and 256b ALUs, and clocking past 5GHz. Apple spends it on scalar resources and deep reordering. I'm not arguing that one tradeoff is universally better than the other, just that they end up similarly big. Since you brought up cache throughput, Zen4 does 32B/cycle between L1 and L2 [1]. Anandtech measured M1's L2 cache throughput at about 440 GB/s across the 4 P-cores [2], which works out to 34B/cycle/core. Which sure sounds like the same per-core throughput to me. > You can just... look at the damn die shot I... did? That's how I estimated L2 cache and tags at 28% of Zen4's 3.84mm^2. Do you believe that that is incorrect, or that the M1 and M2 P-core areas of 2.281-2.756 mm^2 quoted by Semianalysis are incorrect? [1] https://chipsandcheese.com/2022/11/08/amds-zen-4-part-2-memory-subsystem-and-conclusion/ https://chipsandcheese.com/2022/11/08/amds-zen-4-part-2-memo... [2] https://www.anandtech.com/show/17024/apple-m1-max-performance-review/2 https://www.anandtech.com/show/17024/apple-m1-max-performanc...