5 ms·
256 cores on a die. Stunning.
by mrlonglong 9mo ago
256 cores on a die. Stunning.
- jauntywundrkind 9mo agoIntel's Clearwater Forest could be shipping even sooner, 288 cores. https://chipsandcheese.com/p/intels-clearwater-forest-e-core-server https://chipsandcheese.com/p/intels-clearwater-forest-e-core... It's a smaller denser core but still incredibly incredibly promising and so so neat.
- jsheard 9mo agoSomeone needs to try running Crysis on that bad boy using the D3D WARP software rasterizer. No GPU, just an army of CPU cores trying their best. For science.
- zeusk 9mo agoThis has already been tried :) iirc, in the 2016 a quadcore intel cpu ran the original crysis at ~15fps
- mrlonglong 9mo agoAh, I omitted to mention that with 256 cores, you get 512 threads.
- CyberDildonics 9mo ago"E-cores" are not the same
- bri3d 9mo agoThe 32 core / die AMD products are almost certainly Zen 6c, which is the same "idea" as Intel E-Cores albeit way less crappy. https://www.techpowerup.com/forums/threads/amd-zen-6-epyc-venice-introduces-a-radical-package-redesign.344788/#post-5652196 https://www.techpowerup.com/forums/threads/amd-zen-6-epyc-ve... EDIT: actually, now that I think about it some more, my characterization of Zen-C cores as the same "idea" as Intel E-cores was pretty unfair too; they do serve the same market idea but the implementation is so much less silly that it's a bit daft to compare them. Intel E-Cores have different IPC, different tuning characteristics, and different feature support (ie, they are usually a different uarch) which makes them really annoying to deal with. Zen C cores are usually the same cores with less cache and sometimes fewer or narrower ports depending on the specific configuration.
- eigenform 9mo agoie. marketed as "dense" instead of "efficient"
- deleted 9mo ago[deleted]
- eigenspace 9mo agoI was about to reply with an "well, actually..." comment and then I saw that you beat me to it with your edit. Fully agreed, they may be targetting a similar goal, but the execution is so different, and a Intel screwed up the idea so bad that it can really mislead people into assuming that dense Zen cores are the same junk as a Intel E-cores.
- hypercube33 9mo agoI may be wrong, as I'm not an expert but Intel E-Cores are basically decendents of Intel ATOM (I personally really liked the idea of ATOM it was just so nerfed by Intel with its memory limits and platforms) and P cores are derived from the i-Series - two totally different cores. Yes they support the same instruction sets generally, but they are different cores. AMD's approach was to basically trim the fat on Zen as far as they possibly could but keep the core fast and efficient and you end up with the C-Cores. In practice (where I generally live with expertise) AMD's approach lets you move software between cores and it generally doesn't care or know, whereas Intel's method applications definetly do care and can have issues. To me, the difference is moving a virtual machine between two similar CPUs and not having to reboot it (AMDs approach) and having to reboot for one reason or another (Compatibility) -- Intels method.
- bri3d 9mo agoYes, you're right and that's what I discuss in the last paragraph of my comment. And, yes, E-cores are rough descendants of some processors that were sometimes called Atom. Using Intel marketing names is fraught with peril, though (see, Celeron) as "Atom" has referred to many conceptually different microarchitectures over time; the modern E-Core is of no relation to the original in-order "Atom" processor many remember. I've had the same experiences as you with Intel mixed-core desktop parts. They're incredibly difficult to optimize for due to the heterogeneous core mixture, whereas AMD mobile parts are generally more reasonable (you're on a slow core or a fast core, basically), and AMD never made a mixed-core desktop part. However, Intel server parts several years ago switched to E-core only or P-core only, so all of the heterogeneous core mixture issues aren't a thing - you basically have two separate processor generations being sold at once, which isn't particularly surprising or uncommon. With AMD server processor families (linked in my comment), depending on the part's density you get either "slow" or "fast" cores and either "wide" or "narrow" units, so you do still have to think about things a little bit there too. Where Intel really screwed up in general, microarchitecture differences aside, is AVX512. That's the wrench that prevents the same compiled code from running across most Intel parts - they just couldn't decide what they wanted to do with it, whereas AMD just chose to support it and stick with it, even though the throughput for the wide instructions is wildly different between processors.
- tester756 9mo agoBy what logic?
- eigenspace 9mo agoIntel E-cores are basically a different microarchitecture. They often support different instruction sets than their P-cores, have different "instructions-per-clock" rates (IPC), and all sorts of other major differences. They're just very different things, and those differences are responsible for most of the bad reputation that E-cores have. AMD's dense-cores are the same microarchitecture, have the same IPC, use all the same instruction sets. The only real difference between them and regular AMD cores is that their dense cores have less cache, and lower peak clocks.
- tester756 9mo ago>They often support different instruction sets than their P-cores Do they? I thought it caused very significant problems (when there's switch between E and P core) and they avoided it But I cannot find anything about it
- wmf 9mo agoThe P and E cores support different instructions and Intel "fixed" it by disabling instructions on the P-cores. So now they have the same instructions but at the cost of a bunch of wasted silicon.
- eigenspace 9mo ago> Do they? Yes?
- adrian_b 9mo agoThe Intel server CPUs with P-cores support AVX-512, like all current AMD CPUs, and they also support a few extensions not currently supported by AMD, like AMX (Zen 6 will add FP16 arithmetic support in AVX-512, reducing the differences vs. Intel P-core servers). The Intel server CPUs with E-cores, both the current Sierra Forest CPUs with Crestmont cores and the future Clearwater Forest CPUs with Darkmont cores do not support AVX-512 and they are almost identical with the E-cores from Intel laptop/desktop CPUs. Therefore, for demanding applications you cannot run the same programs on Intel servers with P-cores or E-cores, unless they use dynamical dispatch to select at run-time between AVX and AVX-512 libraries, as the gain from AVX-512 can be very substantial and on server applications not using it would lose money by lowering the throughput. The Intel Darkmont cores of Panther Lake and Clearwater Forest are almost identical with the Skymont cores of Arrow Lake and Lunar Lake (the main difference is that the Skymont cores are made by TSMC, while the Darkmont cores are made by Intel in their new 18A CMOS process) and they are extremely similar in die size and in performance with the ARM Neoverse V3 cores from the newly launched AWS Graviton5 (which are known as Cortex-X4 in their smartphone variant). Intel has said that they will eliminate this ISA difference between E-cores and P-cores, but a couple of years might pass until this will reach their server CPUs.
- bee_rider 9mo agoI wonder what Ampere (mentioned in that article) is going to do. At this rate they’ll need to release a 1000 cpu chip just to be noticeably “different.”
- wmf 9mo agoUnfortunately Ampere has fallen pretty far behind AMD. I don't see much point to their recent CPUs.
- fc417fc802 9mo agoAt some point won't the bandwidth requirements exceed the number of pins you can fit within the available package area? Presumably you'll end up back at a low maximum memory high bandwidth GPU design. I wonder how many of these you could cram into 1U? And what the maximum next gen kW/U figure looks like.
- epistasis 9mo agoThat's going to run Cities Skylines 2 ~~really really well~~ as well as it can be run.
- mort96 9mo agoDoes it actually scale well to that many cores? If so, that's quite impressive; most video game simulations of that kind benefits more from few fast cores since parallelizing simulations well is difficult
- Neywiny 9mo agoNo, see https://m.youtube.com/watch?v=44KP0vp2Wvg https://m.youtube.com/watch?v=44KP0vp2Wvg . You're right it didn't scale that well
- epistasis 9mo agoLooks like it may be capped at 32 cores in that video, if they are hitting 25%-30% of a 96 core CPU? Here's analysis of a prior LTT video showing 1/3 of cores at 100%, 1/3 of cores at 50%, and 1/3 idle cores: https://www.youtube.com/watch?v=XqSCRZJl7S0 https://www.youtube.com/watch?v=XqSCRZJl7S0 In any case, CS2 can take advantage of far more cores than most games.
- markhahn 9mo agothese big high-core systems do scale, really well, on the workloads they're intended for. not games, desktops, web/db servers, lightweight stuff like that. but scientific, engineering - simulations and the like, they fly! enough that the HPC world still tends to use dual-socket servers. maybe less so for AI, where at least in the past, you'd only need a few cores per hefty GPU - possibly K/V stuff is giving CPUs more to do...
- rbanffy 9mo ago> not games, desktops, web/db servers, lightweight stuff like that. Things like games, desktops, browsers, and such were designed for computers with a handful of cores, but the core count will only go up on these devices - a very pedestrian desktop these days has more than 8 cores. If you want to make software that’ll run well enough 10 years from now, you’d better start using computers from 10 years from now. A 256 core chip might be just that.
- Neywiny 9mo ago32 cores on a die, 256 on a package. Still stunning though
- bee_rider 9mo agoHow do people use these things? Map MPI ranks to dies, instead of compute nodes?
- wmf 9mo agoYeah, there's an option to configure one NUMA node per CCD that can speed up some apps.
- markhahn 9mo agoMPI is fine, but have you heard of threads?
- bee_rider 9mo agoSure, the conventional way of doing things is OpenMP on a node and MPI across nodes, but * It just seems like a lot of threads to wrangle without some hierarchy. Nested OpenMP is also possible… * I’m wondering if explicit communication is better from one die to another in this sort of system.
- fc417fc802 9mo agoWith 2 IO dies aren't there effectively 2 meta NUMA nodes with 4 leaf nodes each? Or am I off base there? The above doesn't even consider the possibility of multi-CPU systems. I suspect the existing programming models are quickly going to become insufficient for modeling these systems. I also find myself wondering how atomic instruction performance will fare on these. GPU ISA and memory model on CPU when?
- DiabloD3 9mo agoIf you query the NUMA layout tree, you have two sibling hw threads per core, then a cluster of 8 or 12 actual cores per die (up to 4 or 8 dies per socket), then the individual sockets (up to 2 sockets per machine). Before 8 cores per die (introduced in Zen 3, and retained in 4, 5 and 6), the Zen 1/+ and 2 series this would have been two sets of four cores instead of one set of eight (and a split L3 instead of a unified one). I can't remember if the split-CCX had its own NUMA layer in the tree or not, or if they were just iterated in pairs.
- m4rtink 9mo ago640 cores should be enough for anyone
- jsheard 9mo agoTell that to Nvidia, Blackwell is already up to 752 cores (each with 32-lane SIMD).
- fooblaster 9mo agob200 is 148 sms, so no
- jsheard 9mo agoEach SM cluster contains 4 independent 32-wide compute units, and GB202 has 192 SMs, although only 188 of them are enabled on the largest shipping SKU. IMO that makes for 752 "cores", but depending on where you draw the line it could be 188, 752, or 24064.
- fooblaster 9mo agosms is the Nvidia definition of processor, and cuda device properties returns it, not anything else. If you want a marketing number, use cuda cores, it doesn't consistently match to anything in the hardware design.
- markhahn 9mo agono, you really can't. NVidia's use of "cores" is simply wrong. unless you think a core is a simple scalar ALU. but cores haven't been like that for decades. or would you like to count cores in a current AMD or Intel CPU? each "core" has half a dozen ALUs/FP pipes, and don't forget to multiply by SIMD width.
- phkahler 9mo ago640K cores should be enough for everyone.