Area. Golden Cove (12-series) and Raptor Cove (13-series, except for some of the lower SKUs which get rebranded Golden Cove) are obscenely massive. It is something close to 2x the area of zen3 per core, which is not even on a 5nm tier node like Intel 7! and logic density went up 1.8x between TSMC N7 and N5, so this means something like 3.2x the transistor count at the high end. Achieved shrink will be a bit lower, but let's say 3x the transistor count of zen3.
And probably this tends to understate things because that's Golden not Raptor cove. Intel went obscenely huge on caches with Raptor Cove too - they did the same thing as NVIDIA with Ada, and dumped an assload of L1/L2 cache on it too. I don't know off the top of my head but let's say 10-20% bigger for Raptor cores.
https://www.reddit.com/r/hardware/comments/qlcptr/m1_pro_10core_soc_pips_m1_max_to_head_passmarks/hj6gmb3/ https://www.reddit.com/r/hardware/comments/qlcptr/m1_pro_10c...
In contrast Gracemont is much smaller - it is not quite "4x" as advertised, 4 is actually the number of cores in a Gracemont CCX/cluster, but the cluster is somewhat bigger than a Golden Cove core. So the actual core area is 3.26 Gracemont cores per Golden Cove core, and again, Raptor Cove is significantly bigger.
--
So the tradeoff is like - they could have done a 12P0E or something like that, for about the same area as a 8P12E. Which would still lose to a 16P0E Zen3 in multithreaded workloads, most likely.
That's the game Intel is playing - 8P is generally enough for games, but it's area-inefficient to keep scaling like that. But you have these other bulk tasks that just like tons of cores and don't care about peak performance, so, you have a mix of both. The E-cores give you more perf/area and the p-cores give you more peak perf for games/etc. So theoretically it's the best of both worlds, it's not as slow as a full e-core chip would be but it has a lot more MT performance than an all-P-core would.
Unspoken underlying problem being that Intel's P-cores are much, much, much bigger than the competition. Hence they have a much greater need to come up with a "compact" alternative than AMD does. Using that 1.8x logic scaling factor (which is optimistic), a Zen3-on-5nm design would be 1.72mm2 which is just about the same size as Gracemont. So Intel's "e-core" is about as big as AMD's p-core! Hence why they are much more focused on a whole new core design, where AMD just densifies the existing one (high efficiency/high-density libraries, reduced cache, back to 4-core CCX, etc). Squeeze that last 30% and call it a day.
On a more tactical level, I think it also is a move to force people to use Gracemont and start writing code for it. Long-term, your P-core being 3x the size of your competitors' is not sustainable and they need to pivot away from the existing P-core design (lakes/coves), it obviously is just a mess internally from 3 decades of tech-debt. Nobody really cares about the atom chips, despite them being pretty good for a long time now (my J5005 NUC made a great thin client during the pandemic, I use them for HTPCs, etc). Well, now you have to care, or you're leaving performance on the table on the mainstream intel chips. It's not just "intel loves big.little" or "needs big.little for area" but also "big.little" is a way for them to start getting the "little" cores into running real-world code, because in the long term they need to kill the coves off (and perhaps replace them with a mont-derived alternative).
(my suspicion is that this is a case of Conway's Law in action, and the architecture of the Lake/Cove family resemble the Intel organizational chart, and since Intel is a giant knot, that's the processor architecture they produce, and they've been doing that for at least 20 years. In hindsight Pentium 4 was the warning sign of the internal rot, they got it back together for a while but after the sandy bridge era they collapsed and everything since then is probably just more and more tech debt and kludges stacked on.)
--
Also, frankly, the e-core's "CCX" design makes sense. Tiering your interconnect/cache is what AMD has done very successfully - you have 2 CCXs per CCD (on zen2), 8 CCDs per socket. And that lets them decompose the interconnects into manageable pieces - 4 cores per CCX is a simple all-connected topology. Those talk to 4 quadrants on the IO die, which is a simple topology. If you want to talk to the other CCX, you have to go through the quadrant/IO die, so there is no "special case" there. It's all just a composition of simple pieces.
Ringbuses get annoying/inefficient past about 10-12 cores, which is why Intel abandoned them in server after broadwell-EP (with its "dual ring" design). But a mesh of individual cores also has this huge latency penalty, and consumes a bunch more area, and (in practical configurations) still tends to be very bottlenecked unless you spend an even higher amount of area on it.
What's the middle-ground? You group the cores into clusters/CCXs and you either have a mesh-of-CCX or a ring-of-CCX or some other tiered structure. And you can break the "tile" idea down into tiers too - a tile is a ring or mesh of cores, and then you have a mesh of tiles, but these are separate logical tiers and don't need to interact.
It is the usual HPC networking problem - connecting 1,2, or 4 nodes is easy, with simple all-connected or hypercube topologies, with a small number of links. A hypercube requires only 2 links per node for 4 nodes. An all-connected topology requires only 3 links. And you can solve for modestly higher numbers with something like a ringbus (which gets a lot of flak but it's an extremely performant network structure, and AMD uses them too for their 8-core CCX). But that falls apart with higher numbers of nodes too, and big switched-fabric networking chips or backbone switches are some of the largest and most expensive chips manufactured, a 32-port 400gbe switch (idk, whatever) is gonna be a big beefy boy in itself, that type of thing often hits 750mm2+ of silicon on the latest nodes.
You need something that both scales in terms of network hardware/area/power, and also performs in terms of actual latency and throughput. That's super difficult (and the best topologies are de-facto "tiered" anyway like hypercube or butterfly), so the best strategy is to introduce this tiering. And AMD has meticulously stayed in the limit of "the IO die is a simple hypercube topology of quadrants" and "the CCX is a 4C all-connected or a 8C ringbus", and just composed these simple things together with tiering.
I think low-key Gracemont is important because it's Intel tinkering with the same concept - and they're doing mesh-of-tiles with sapphire rapids too (not sure what topology is inside each tile though). Because they can't have 14+ stops on the ringbus (memory controller, 12 cores, iGPU, etc) and the purist "mesh of single cores" topology obviously didn't work with skylake-SP.
https://www.anandtech.com/show/10158/the-intel-xeon-e5-v4-review/2 https://www.anandtech.com/show/10158/the-intel-xeon-e5-v4-re...
https://www.anandtech.com/show/11544/intel-skylake-ep-vs-amd-epyc-7000-cpu-battle-of-the-decade/5 https://www.anandtech.com/show/11544/intel-skylake-ep-vs-amd...
https://www.anandtech.com/show/14694/amd-rome-epyc-2nd-gen/2 https://www.anandtech.com/show/14694/amd-rome-epyc-2nd-gen/2
https://www.anandtech.com/show/16529/amd-epyc-milan-review/4 https://www.anandtech.com/show/16529/amd-epyc-milan-review/4
--
Anyway, I wish they would do all-P-core too, and theoretically that exists, it's Sapphire Rapids, and there is a workstation/HEDT variant, it's just super expensive and has massive power transient problems (you need 500W of headroom, literally, you will crash if you don't have a 1kw+ PSU, they are not kidding about 1300W being the recommended) that might be sending them back to the drawing board for another stepping. And there is also all-e-core chips too, that's Sierra Forest... but it seems like a tentpole customer pulled out (rumored to be facebook iirc) because Bergamo, AMD's compact-core based Epyc server chip, is more attractive. And so they have reduced the scope of Sierra Forest, it now tops out at 2 of the medium chiplets and the big chiplets are canceled entirely (where they planned to use up to 4 of the big ones).
MLID is such an unreliable source that I hesitate to recommend him, I like his content and listen to him a lot, but you really need to understand the broader context of the market/etc to know whether what he's saying makes sense. But he does tend to have some interesting guests who are usually way better than he is, and one of his recent guests was a boutique PC builder who specializes in digital audio workstations (which need to be super low latency/etc). And they talk about Sapphire Rapids workstation and some of the things being discussed around it.
https://www.youtube.com/watch?v=_HJu5xt43iQ&t=3603s https://www.youtube.com/watch?v=_HJu5xt43iQ&t=3603s (and the previous segment too)
Sierra forest discussion: https://www.youtube.com/watch?v=QlTZCDEFUFg&t=4200s https://www.youtube.com/watch?v=QlTZCDEFUFg&t=4200s
General intel discussion: https://www.youtube.com/watch?v=BNXlRdAKWTE https://www.youtube.com/watch?v=BNXlRdAKWTE
Oh, so, I got lost in the appendix explaining network topology, but to circle back on this: yes, yield, and cost. It would be a very large total area too even if it yielded well. Just too expensive in general.
Alder Lake 8P8E: 215.25mm2
Alder Lake 6P0E: 162.75mm2
Raptor Lake 8P16E: 257mm2
For comparison:
8700K: 149.6mm2
9900K: 174mm2
10900K: 206.1mm2
Zen2 CCD: 74mm2
Zen3 CCD: 83.74mm2
That's actually pretty big for a consumer processor already. And it's all in monolithic 5nm(-tier node), which isn't cheap even if it yields fine. So them having a uarch that's at a pretty bad area disadvantage isn't good, and tbh they obviously aren't delivering on any kind of efficiency promise.
Physics is getting hard and wafer costs are spiraling pretty bad, which is why AMD is exploring advanced packaging/etc. Doesn't always work though - like RDNA3. Data movement still seems to be very expensive, although 2.5d and 3d stacking (and direct-bonding) will mitigate this somewhat. But advanced packaging means moving a lot more data, and you have to be careful of what lives on what side of what links. Cache being on the other side of the infinity links (not infinity fabric!) in RDNA3 seems like potentially a specific problem with the design, since you pay the cost for the data movement to the cache and not just the data movement for the memory.
I think you're right they probably could do it if they wanted etc, maybe sell it as a pseudo-HEDT (especially if you can glue together a pair of dies directly to 2x the normal core count - and if you can glue together 2x16C all-P-core designs that's fine for HEDT for a lot of things imo!). But the price would probably be fairly high (16C would be like, probably $700-900) and the power would still be quite high (intel does not win at any power bracket right now even with limits, it's just less bad if you limit it to 150W), etc. Maybe some of the power stuff would go away if you got rid of the split-brain big/little clusters on a ring thing, but, even if you went with 16 P-cores on a ring, the latency would still go up a lot, and you'd notice it because the stuff you want it for is gaming/etc. The latency would hurt gaming IPC a decent chunk imo, or you'd have to go to a double-ring like broadwell.
It's a mess and this is the point where the ringbus scaling craps out, is my point with the latency discussion. It seems hard to have more than about 8C or 12C per "tier". Even Bergamo (AMD's new e-core variant of Epyc) is 16C of Zen4C per CCD, but it's 2 CCXs of 8. Broadwell dual-ringbus is 2 tiers of 12 cores each. The subsequent Intel chips moved to the mesh. Alder/Raptor do 8P+4 e-core clusters (12 nodes). Etc. You can add more tiers of 8-12, but about 8-12 nodes per tier seems to be the limit that scales well due to interconnect bandwidth/etc, just historically imo. Interesting convergence.
(plus a couple nodes for pcie agents and iGPU and memory controller and shit I'm not counting here, not stops just just cores)
https://www.anandtech.com/show/10158/the-intel-xeon-e5-v4-review/2 https://www.anandtech.com/show/10158/the-intel-xeon-e5-v4-re...
I think strategically they want and need to keep selling the e-cores though, it's not what's right for you, it's what's right for them and their migration path. Some of these pieces it's hard to see how you do everything in a single go - it's tough to go from "everything is the same" to "lol CMT with 3 slow/1 fast thread controlled by this thread director that wants to talk to your OS scheduler". But the theoretical end-state of "big.little within a core cluster" or "within a CMT core" is pretty neat at least, that would mitigate the latency problems of dedicated "little core clusters". And this is one of jim keller's ideas apparently, while he was at intel (briefly, lol)
Now again, to say something nice here: the p-cores are pure out-of-order monsters. Very wide decode units, lots of execution resources, etc. It is the same as the "zen4 vs zen4-X3D" split, stuff that prefers zen4 over x3d also really prefers raptor cove, it's an execution monster and it does it all on just 8 p-cores. it just also uses more energy to do it, and cache can handle some other situations where the working set helps cover some useful working set. And the e-cores do give you a ton of the performance equivalence of having a wider AMD processor in MT workloads (if they're not latency-sensitive, ie cinebench and video encoding), just not at particularly great power compared to AMD's 16 p-cores clocked much lower.
I like my Atom processors a lot (and I've looked seriously at denverton etc). They're not bad cores at all, but Alder/Raptor just put them in the worst possible place, with too much voltage (no DLVR!), meaning you might as well goose the whole thing and go for power etc, and the latency isn't flattering.
I would have loved the idea of Sierra Forest-HEDT with like 192 / 256 cores (or whatever specifics) enabled or whatever, if they could get that to a relevant price for enthusiasts it'd be amazing, the HEDT market is super dead and the e-cores are good enough nowadays. That would be a super high-value place to put a product, if it's not moving adequately in the server market etc. Give me 5820K level value for something the client platform cannot do, and Intel benefits from actually getting a foothold on a market that is willing to tinker and build stuff.
But everything intel does has come so late that it almost doesn't matter, it's a sidegrade at best, at much worse efficiency etc. Just buy a 7700X or 7800X3D or 7950X/X3D. Let alone the slaughter in the server market etc. Sapphire Rapids is not great either, lots o weird power shit, and maybe it needs another stepping to fix it? ok, or, if you're a hyperscaler, amd will ship you zen5 samples probably within the next 3 months, and they can actually execute and deliver.
Hard to see that changing anytime soon either, I don't even see green shoots, I think they're in a death spiral tbh. I see the 2.5gbe nic team failing over and over (I225/I226), I see sapphire rapids taking far too many steppings and base die changes, I see DLVR still not working, I see no plan for AVX-512, I am guessing they have continual problems with integration/packaging, I see every product being a one-off with no reusability, etc. They're too important to let go under but they're in deep shit and they have to do a turnaround on an understaffed underpaid employee force etc, while battling deep internal-culture rot and middle-management warfare etc. It's gonna be a while before they're relevant imo.
I, too, have looked at big.little in 12/13th gen laptops and just ehhh is the world ready for that yet? I thought 11th gen was kinda attractive for that reason, the last all-p-core uarch (oh and you get avx512 too, etc). Orrrr you just buy something with a 7940HS/HX/whatever and get 8 big old Zen4 cores... with AVX-512.... and that's every product segment and it's going to get worse, pretty much. Intel still coasts on massive availability but at a technical level AMD is clowning them in pretty much most enthusiast or enterprise use-cases.
(one exception, usually, is system stability. AMD's AM4 USB dropout glitch isn't really fixed despite lots of effort, neither is the AM4/5 fTPM stutter bug, and a physical TPM header is something to look for on an AMD board lol, because fTPM is broken and causes random stutter (TPM operations getting blocked by some single-threaded UEFI process in vendors' UEFI implementation, is the internet speculation). Intel mostly is better about not having that shit, for now. In the past I have heard of a lot of problems with AMD chips in servers too, just weird linux problems etc, (and of course segfault affected a lot of early Zen1/1000-series chips in a lot of things, the scope was downplayed p. bad), but that's scandalous hearsay and tbh today I think Asrock X570D4U-2T or ROMED8-2T or GENOAD8X-2T owners etc are happy, people would report problems etc. Intel does have the support story of being the default... right up until they won't anymore.)
They have some neat ideas but AMD is gonna leap forward again with Zen5 too, everyone has neat ideas that will be coming to fruition in 3-5 years. Zen4 was the easy stuff - a pretty minor port of zen3 to 5nm, with DDR5, with AVX-512, and some tweaks to open up architectural or timing bottlenecks to push clocks etc, while they did the DDR5 switchover stuff. Some cleanup (and they got it pretty much to 6 GHz in peak 1T lol) but all the interesting stuff is coming next year in zen5, it's gonna be much wider etc (similar to intel's own width increase with golden) etc, it's expected to be a pretty significant increase. this is a big rework of the whole thing to clean up and scale higher. so this 14th-gen stuff will be going up against zen5 being probably 20-30% faster in general performance, it'll be a pretty decent uplift. They're in trouble in pretty much every product segment already, it really seems like they struggle to even get the product out the door these days.
https://en.wikichip.org/wiki/intel/microarchitectures/alder_lake#Die https://en.wikichip.org/wiki/intel/microarchitectures/alder_...
https://en.wikichip.org/wiki/intel/microarchitectures/raptor_lake#Die https://en.wikichip.org/wiki/intel/microarchitectures/raptor...
https://www.techpowerup.com/297506/intel-raptor-lake-core-i9-13900-de-lidded-reveals-a-23-larger-die-than-alder-lake https://www.techpowerup.com/297506/intel-raptor-lake-core-i9...
https://en.wikichip.org/wiki/intel/microarchitectures/coffee_lake#Die https://en.wikichip.org/wiki/intel/microarchitectures/coffee...
https://www.techpowerup.com/267649/intel-core-i9-10900k-der8auer-de-lidding-reveals-accurate-die-size-measurements https://www.techpowerup.com/267649/intel-core-i9-10900k-der8...
https://en.wikichip.org/wiki/amd/microarchitectures/zen_2#Die https://en.wikichip.org/wiki/amd/microarchitectures/zen_2#Di...
https://wccftech.com/amd-ryzen-5000-zen-3-vermeer-undressed-high-res-die-shots-close-ups-pictured-detailed/ https://wccftech.com/amd-ryzen-5000-zen-3-vermeer-undressed-...
https://wccftech.com/amd-epyc-bergamo-cpu-die-detailed-16-zen-4c-vindhya-cores-per-ccd-35-percent-smaller-core-area/ https://wccftech.com/amd-epyc-bergamo-cpu-die-detailed-16-ze...
https://en.wikipedia.org/wiki/List_of_Intel_Core_i7_processors https://en.wikipedia.org/wiki/List_of_Intel_Core_i7_processo...
https://en.wikipedia.org/wiki/Intel_Graphics_Technology https://en.wikipedia.org/wiki/Intel_Graphics_Technology
https://en.wikipedia.org/wiki/Tegra#Models https://en.wikipedia.org/wiki/Tegra#Models
https://en.wikipedia.org/wiki/CUDA#Version_features_and_specifications https://en.wikipedia.org/wiki/CUDA#Version_features_and_spec...
(just wanted to call out: wikichip is a lovely site, lots of randomly useful information there. and wikipedia also has a number of useful lists of cpus/gpus/mobile SOCs/etc with characteristics listed, and good sources for CUDA compute capability etc. Don't sleep on wikipedia as a quick reference for what the boost is on random xeon sku xyz, and so on.)