10 ms·
Just to quantify your adjectives, per the Anandtech article: > The M1 is really wide (8 wide decode) In contrast to x86 CPUs which are 4 wide decode. > It ha
by libria 6y ago
Just to quantify your adjectives, per the Anandtech article:
> The M1 is really wide (8 wide decode)
In contrast to x86 CPUs which are 4 wide decode.
> It has a huge 630 deep reorder buffer
By comparison, Intel Sunny/Willow has 352.
- deleted 6y ago[deleted]
- natchy 6y agoSo Intel and AMD are capable of building a chip like this, but the ambitious size meant it was more economically feasible for Apple to build it themselves? (Not a hardware guy)
- ArchOversight 6y agoRead the article linked, it explains why Intel and AMD are unable to throw more decoders at the problem.
- outworlder 6y agoMaybe they are, assuming there's sufficient area in the die for this. They would likely still be massive power hogs.
- masklinn 6y agoIt's not that it was more economical, but that at least some of these AMD and Intel would not benefit from due to the ISA: x64 instructions can be up to 15 bytes, so just finding 8 instructions to decode would be costly, and I assume Intel and AMD think more so than the gains from more decoders (you couldn't keep them fed enough to be worth it, basically).
- SlipperySlope 6y agoWhen ARM migrated to 64 bits, wisely they left the 32 bit instructions behind. So easy to decode and uniform instruction length. On the other hand, AMD64 can run 8088 16 bit instructions, which means that it must handle a variety of instruction lengths.
- lmilcin 6y agoNeither Intel nor AMD are capable of doing for a very basic reason, there is no market for it. You can't just release a CPU for which there is no operating system. Apple can pull it off because they already own entire stack from hardware to operating system to cloud services and the can swap out a component like CPU for a different architecture and release new version of OS that supports it. Apple, by creating new CPU, replace a part of the stack that is owned by Intel by their own which only strengthens them position even if it did not improve any performance. Apple is invulnerable to other companies copying the CPU and creating their own because they are not really a competition here. Apple sells an integrated product of which CPU is just one component.
- throwaway_pdp09 6y ago> You can't just release a CPU for which there is no operating system sure you can. That's what compilers are for.
- loeg 6y agoGP probably means that you won't be able to sell it, even if there is a compiler. (Not true in the super embedded space, sure.)
- skavi 6y agoIntel had such an attitude once before.
- throwaway_pdp09 6y agoDonald Knuth said "The Itanium approach...was supposed to be so terrific—until it turned out that the wished-for compilers were basically impossible to write."[82] https://en.wikipedia.org/wiki/Itanic https://en.wikipedia.org/wiki/Itanic So they didn't have the needed compiler
- pjmlp 6y agoI wonder how HP and Microsoft managed to port HP-UX and Windows without a compiler.
- skavi 6y agoNot necessarily. Samsung used to make custom cores that were just as large if not larger than Apple’s (amusingly the first of these was called M1). Unfortunately, Samsung’s cores always performed worse and used significantly more power than the contemporary Apple cores. Apple’s chip team has proven capable of making the most of their transistor budget, and there’s reason to believe neither Intel nor AMD could achieve Apple’s efficiency even if they had the same process, ISA, and area to work with.
- The_Colonel 6y ago> and there’s reason to believe neither Intel nor AMD could achieve Apple’s efficiency even if they had the same process, ISA, and area to work with. What's that reason?
- Accujack 6y agoI think it's "faith".
- yazaddaruvala 6y agoFaith implies no data. Why will Apple always out compete Intel and other non-vertically-integrated systems? Margins, future potential, customer relationship and compounding growth/virtuous cycle. The margins logic is simple, iPhones and MacBooks make tons more money per unit compared with a CPU. Imagine if improving the performance of a CPU by 15% makes the demand increase by 1%. For Apple improving the performance of a CPU by 15% makes the demand increase by 1% for the whole iPhone or MacBook. For this reason alone, Apple can invest 2-5x more R&D into their chips than everyone else. The future potential logic is more nuanced: 1. Intel's/whoever's 10 year vision is to build a better CPU/GPU/RAM/Screen/Camera because their customers are the companies buying CPUs/GPUs/Screens/Cameras/RAM. They are focused on the metrics the market has previously used to measure success and want to build to optimize for those metrics e.g. performance per dollar. Intel doesn't pay for the electricity in the datacenter nor through its customers' complaints about battery life. RAM manufacturers aren't looking at Apple's products and asking, "do consumers even replace still RAM?" i.e. they are focused on "micro"-trends. 2. Apple's vision is to build the best product for customers. They look at "macro"-trends into the future and apply their personal preferences at scale. For example, do people even still need replaceable RAM? Will they want 5G in the future, or can we improve the technology to replace it with direct connections to a LEO satellite cluster? The customer relationship logic: Lets take one such example of a macro-trend, VR and other wearables. Apple is tracking these trends and can "bet on" them because its in full control but Nvidia, Intel, etc typically don't want to "bet on" these numbers because even if they are fully invested, their partners (which sell to consumers) might back out. Apple also isn't "betting on" because it has a healthy group of early adopters that trust Apple and will buy and try it even tho a "better" product in the same market segment isn't purchased. Creating/retaining that customer relationship lets Apple over invest into keeping heat (i.e. power) low because its thinking about the whole market segment that Apple's VR headset can start to compete in and collect more revenue from. Compounding growth/virtuous cycle logic is also relatively simple: Improving the metrics in any of these 3 previous pillars manipulatively improves the other pillars. i.e. better customer relationship increases cashflow, increseses R&D funding, 1. improves product, improving customer relationship; or 2. reduces costs, increasing margins, and loops back to increasing cash flow.
- amluto 6y agox86 instructions are variable length with byte granularity, and the length isn’t known until you’ve mostly decoded an instruction. So, to decode 4 instructions in parallel, AIUI you end up simultaneously decoding at maybe 20 byte offsets and then discarding all the offsets that turn out to be in the middle of an instruction. So the Intel and AMD decoders may well be bigger and more power hungry than Apple’s.
- WhyNotHugo 6y agoThe problem is the market. Windows only a single architecture, so they can't really deviate from that. Sure, windows can switch (or, apparently, run on ARM), but due to the fact that windows applications are generally distributed as binaries, lots of apps wouldn't work. Linux users would have far less issues, and would be a great clientele for a chip like this, but probably too niche a market, sadly.
- kjs3 6y agoWindows only a single architecture People forget that at launch Windows NT ran on MIPS & DEC Alpha in addition to x86. The binary app issue was a killer for the alternative archs.
- monocasa 6y agoDEC Alpha NT could run X86 code thanks to FX!32, and faster than a core you could buy from Intel at the time.
- pjmlp 6y agoIndeed, but it didn't had anything that justified actually paying big bucks for an Alpha.
- foobiekr 6y agowell, for some things. fx32 for the apps people wanted though was deficient. The NT3.1-era Alphas didn't have byte-level performance so things like Excel, Word, etc. all ran terribly, as did Emacs and X. I supported a lab of Alphas running Ultrix and they were dogs for anything interactive and fantastic for anything that was a floating point application.
- kjs3 6y agoYeah...anyone who thinks fx32 was faster in the real world than a native Intel core never actually ran it.
- 6y ago
- redraga 6y agoI can't comment on the economics of it but I can comment on the technical difficulties. The issue for x86 cores is keeping the ROB fed with instructions - no point in building a huge OoO if you can't keep it fed with instructions. Keeping the ROB full falls on the engineering of the front-end, and here is where CISC v RISC plays a role. The variable length of x86 has implications beyond decode. The BTB design becomes simpler with a RISC ISA since a branch can only lie in certain chunks in a fetched instruction cache line in a RISC design (not so in CISC). RISC also makes other aspects of BPU design simpler - but I digress. Bottom line, Intel and AMD might not have a large ROB due to inherent differences in the front-end which prevent larger size ROBs from being fed with instructions. (Note that CISC definitely does have it's advantages - especially in large code foot-print server workloads where the dense packing of instructions help - but it might be hindered in typical desktop workloads) Source: I've worked in front-end CPU micro-architecture research for ~5 years
- hajile 6y agoHow do you feel about RISC-V compact instructions? The resulting code seems to be 10-15% smaller than x86 in practice (25-30% smaller than aarch64) while not requiring the weirdness and mode-switching associated with thumb or MIPS16e. Has there actually been much research into increasing instruction density without significantly complicating decode? Given the move toward wide decoders, has there been any work on the idea of using fixed-size instruction blocks and huffman encoding?
- redraga 6y agoI can't really comment on the tradeoffs between specific ISAs since I've mainly worked on micro-arch research (which is ISA agnostic for most of the pipeline). As for the questions on research into looking at decode complexity v instruction density tradeoff - I'm not aware of any recent work but you've got me excited to go dig up some papers now. I suspect any work done would be fairly old - back in the days when ISA research was active. Similar to compiler front-end work (think lex, yacc, grammar etc..) ISA research is not an active area currently. But maybe it's time to revisit it? Also, I'm not sure if Huffman encoding is applicable to a fixed-size ISA. Wouldn't it be applicable only in a variable size ISA where you devote smaller size encoding to more frequent instructions?
- deleted 6y ago[deleted]
- trynumber9 6y agoBut in one x86 instruction you often have more complex operations. Isn't that part of the reason why Sunny Cove has only 4 wide decode but still the decoders can yield 6 micro-ops per cycle? That single stat makes it look worse than it is in reality, I think.
- bob1029 6y agoThe whole principle of CISC (v RISC) is that you have more information density in your instruction stream. This means that each register, cache, decode unit, etc. is more effective per unit area & time. Presumably, this is how the x86 chips have been keeping up with fewer elements in terms of absolute # of instructions optimized for. The obvious trade-off being the decode complexity and all the extra area that requires. One may argue that this is a worthwhile trade-off, considering the aggregate die layout (i.e. one big complicated area vs thousands of distributed & semi-complicated areas) and economics of semiconductor manufacturing (defect density wrt aggregate die size).
- zozbot234 6y agoExcept that RISC-V ISA manages to reach infornation density on par with x86 via a simple, backwards-compatible instruction compression scheme. It eats up a lot of coding space, but they've managed to make it work quite nicely. ARM64 has nothing like that, even the old Thumb mode is dead.
- pclmulqdq 6y agoZen 2 has has 8-wide issue in many places, and Ice Lake moves up to 6-wide. Intel/AMD have had 4-wide decode and issue width for 10 years and I'm glad they're moving to wider machines. Edited "decode" to "issue" for clarity.
- BlackFingolfin 6y agoCould you explain what you mean with "8-wide decode in many places" ? How is that possible, isn't instruction coding kinda always the same? I.e. always 4-wide or always 8-wide, but not sometimes this and sometimes that. All sources I could find say it is 4-wide, so I'd also be interested if you could perhaps give a link to a source?
- coliveira 6y agoI guess this requires extending the architecture to 8-wide instructions when it makes sense.
- BlackFingolfin 6y agoWhat do you mean with "8-wide instructions", and what does that have to do with multiple decoders?
- jlouis 6y agoX86 isn't fixed width instructions. Depending on the mix you may be able to decode more instructions. And if you target common instructions, you can get a lot of benefit in real world programs. Arm is different but probably easier to decode. So you can widen the decoder.
- BlackFingolfin 6y agoOK, I could see how one could implement a variable width instruction decoder (e.g. "if there are 8 one-byte instructions in a row, handle them, otherwise fallback to 4-way decoding" -- of course much more sophisticated approach could be made). But is this actually done? I honestly would be interested in a source for that; I just searched again and could find no source supporting this (but of course I may have simply not used the right search, I would not be surprised by that in the least). E.g. https://www.agner.org/optimize/microarchitecture.pdf#page216 https://www.agner.org/optimize/microarchitecture.pdf#page216 makes no mention of this and calls AMD Zen (version 1; it doesn't saying anything on Zen 2/3). I did find various sources which talk about how many instructions / µops can be scheduled at a time, and there it may be 8-way, but that's a completely different metric, isn't it?