4 ms·
4 x86 decoders, all modern x86 CPUs have micro-op caches that are wider than 4 and the retire width of the machines are 6 or greater. Saying it can process twic
by syberslidder 6y ago
4 x86 decoders, all modern x86 CPUs have micro-op caches that are wider than 4 and the retire width of the machines are 6 or greater. Saying it can process twice as many instructions per clock cycle is an incredibly incorrect statement.
- syberslidder 6y agoTo follow-up on my own point, once an x86 instruction has been decoded into a micro-op, it gets stored in a micro-op cache, where the vast majority of the frontend spends it's time fetching from. Additionally, "in flight" is a function of frontend width, Reorder buffer size, additional Out of order structures sizes and throughput depends on those plus the retire width. Additionally, IPC alone is a poor metric for performance analysis because run time is a factor of IPC and clock speed. M1 has higher IPC because it runs at around 3 GHz compared to x86 CPUs running at 4 or 5 GHz. Apple optimizes for IPC, they are all different tradeoffs. Lastly, Apple's cores are probably the biggest in the industry. Compare their new generation of cores, they had to jumbo-size every structure (much larger than most x86 cores). So while it is impressive from a design point, from a micro-architecture point their cores have ballooned in size and are on 5 nm. Their architects are very good, but not godly like I see many tech sites ascribe them to be. Experience: Designed SPARC, ARM, and x86 CPUs
- nielsbot 6y agoHow does this square with their relatively low power consumption? (Honestly asking--I don't know about this stuff)
- userbinator 6y agoI suspect that's largely due to the smaller process size.
- ben-schaaf 6y agoTo add to this: TSMC states either 15% performance or 30% efficiency gains on 5nm (relative to 7nm). Applying the 30% efficiency gains to the similarly performing Zen 3 CPUs gets you to similar consumption as the M1 (in a load scenario, big little has idle advantages).
- GeekyBear 6y agoThe theory when the first Apple 5nm chip was announced is that Apple took the power efficiency gain. >The one explanation and theory I have is that Apple might have finally pulled back on their excessive peak power draw at the maximum performance states of the CPUs and GPUs, and thus peak performance wouldn’t have seen such a large jump this generation, but favour more sustainable thermal figures. Apple’s A12 and A13 chips were large performance upgrades both on the side of the CPU and GPU, however one criticism I had made of the company’s designs is that they both increased the power draw beyond what was usually sustainable in a mobile thermal envelope. This meant that while the designs had amazing peak performance figures, the chips were unable to sustain them for prolonged periods beyond 2-3 minutes. Keeping that in mind, the devices throttled to performance levels that were still ahead of the competition, leaving Apple in a leadership position in terms of efficiency. https://www.anandtech.com/show/16088/apple-announces-5nm-a14-soc-meagre-upgrades-or-less-power-hungry https://www.anandtech.com/show/16088/apple-announces-5nm-a14...
- syberslidder 6y agoSimilar to overall performance, power consumption has many factors. As others have pointed out, a new process is a big part of it. Additionally, they have very efficient and well designed small cores in their SoC and one benefit their full vertical integration is their OS has very good schedulers for maximizing time on small power-efficient cores without impacting performance in a noticeable way. At a core level, I assume they have very good clock and power gating, so even if the cores have these massive structures, they might be able to power down some or all of them. For instance, adding a 4th vector unit might seem expensive, but most workloads are not running vector instructions, so all 4 of those units can be powered off for the time slice that application is running. Apple also has very good physical design teams that produce custom macros for almost the entire chip, which individually are very minor but do add up significantly.
- yxhuvud 6y agoNot necessarily. The problem is that it doesn't mean anything because they have different instruction sets. That makes comparing ipc pretty useless.
- alwillis 6y agoSaying it can process twice as many instructions per clock cycle is an incredibly incorrect statement. It's not as incredibly incorrect as you may think. An x86 instruction can be as big as 15 bytes and there's no easy way for the decoder to know where one instruction ends and the next one begins. All ARM instructions are one size, making instruction decoding more efficient and makes out of order processing faster as well. More details at https://debugger.medium.com/why-is-apples-m1-chip-so-fast-3262b158cba2 https://debugger.medium.com/why-is-apples-m1-chip-so-fast-32...
- jeffbee 6y agoYou keep posting this but it doesn't become more true just by posting it over and over again.
- reitzensteinm 6y agoParent's comment was about micro-op caches. Modern x86 processors in a relatively tight loop do not decode the instructions each iteration. I suspect you haven't internalized this, because reiterating the complexity of decoding instructions just isn't a valid response. The entire point is avoiding that cost.