5 ms·
> space heaters posing as x86 laptops The instruction set has nothing to do with this. Apple has better numbers because it uses the most advanced manufacturing
by sprash 5y ago
> space heaters posing as x86 laptops
The instruction set has nothing to do with this. Apple has better numbers because it uses the most advanced manufacturing process.
- marcan_42 5y agoBut it does. x86 is hell to decode in parallel, because it's variable length (and in a complex way), so Intel can't build CPUs that keep a wide pipeline fed with instructions efficiently like Apple can, with the 8-wide decoder in the M1. That's one of the things that makes the M1 special, and how it gets away with lower clocks and power consumption than competing x86 designs.
- jabl 5y agoCurrent x86 as well as many Arm cores have micro-op caches, so they don't need as wide decoders as a design that doesn't have such a thing, like M1. (That doesn't take away from the fact that M1 is a very impressive design, of course)
- userbinator 5y agoVariable length doesn't matter that much, you just have something that has a wide shifter on the front. And instruction density is important because it saves cache. A single x86 instruction can take the place of several RISC ones, which is why they needed to have it decode so wide. The M1 is almost entirely a process advantage.
- marcan_42 5y agoIt's not about a shifter, it's about determining instruction boundaries. With variable-length instructions the next instruction start depends on the previous one, making it an inherently serial process. Working around that to decode in parallel is not easy and gets superlinearly more complex the wider you make the decoder. A fixed instruction size architecture like ARM64 doesn't have to deal with any of that.
- userbinator 5y agoA fixed instruction size architecture like ARM64 doesn't have to deal with any of that. What it does have to deal with is needing several times higher fetch bandwidth. It's all a bunch of tradeoffs. Related article: https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-matter/ https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
- Dylan16807 5y ago> It's all a bunch of tradeoffs. Fixed vs. variable is a bunch of powerful tradeoffs. The way x86 can stack a bunch of prefixes on an instruction is pretty bad though. You have to interpret many pieces of an instruction to find the size.
- paulmd 5y agoYou can certainly argue that's because of other design choices (optimizing for high IPC at low clocks by trading off space/etc) but it's certainly not just because of their manufacturing process. Both the A15/M2 architecture and AMD Zen4 will be on TSMC N5P this year, it's pretty doubtful that AMD will be able to catch Apple in IPC or perf/watt imo. And AMD is currently the more "low clock/high IPC" of the two major x86 vendors, so they're the easiest comparison, Intel is really clock-focused still (see: today's Anandtech review comparing AMD 6000/Zen3+ against Alder Lake Mobile and note the power scaling numbers). https://www.anandtech.com/show/17276/amd-ryzen-9-6900hs-rembrandt-benchmark-zen3-plus-scaling/9 https://www.anandtech.com/show/17276/amd-ryzen-9-6900hs-remb... Anyway Jim Keller's "I'm sure x86 isn't dead yet" aside, it seems undeniable that tweaking the instruction set to enable deeper reorder and scale the frontend wider should have performance benefits. Saying out-of-order depth doesn't matter is like saying that code density doesn't matter, or speculation depth doesn't matter. These are things that can be simulated, or measured on real-world code, I don't have numbers at-hand but it seems self-evident that there are metrics that those tweaks should improve on. It doesn't mean x86 is at the end of the line, but if you told me that a 50-year legacy ISA (even if it's been cleaned up a lot over the last 30) had maybe a 10-20% performance, perf/watt, or area/watt benefit due to "intentional design" (aka tuning to the task) and the knowledge of hindsight - I have no reason to doubt that as being facially untrue. 10-20% is still "competitive", so Keller isn't wrong, but it also still puts x86 at a disadvantage in the long term. That's a generation of memory (Zen3+ got 12% moving to DDR5) or most of an architectural generation or maybe a half-node step that x86 would have to stay ahead to maintain parity (not just competitiveness).
- jabl 5y agoThis article suggests that on a Haswell, decoders consumed between 3 and 10% of the power, depending on the workload tested: https://www.usenix.org/system/files/conference/cooldc16/cooldc16-paper-hirki.pdf https://www.usenix.org/system/files/conference/cooldc16/cool... On newer wider and deeper designs this number is most likely smaller, and of course the decoder on Arm or RISC-V consume more than 0% too. So most likely, all in all, the "x86 tax" in terms of power consumption is in the low single digit %.
- cultofmetatron 5y ago> Apple has better numbers because it uses the most advanced manufacturing process. on the contrary, arm is a simpler isa to implement which means radically less circuitry for reading the instruction. modern Intel cpus go so far as having entire systems that translate external x86 instructions to an internal risc based one. That's a huge source of wasted power ie: heat.
- kllrnohj 5y agoARM also decodes to uOps, and no that decoding is not a huge source of wasted power. The vast, vast majority of power (and thus heat) in x86 CPUs (or ARM for that matter) is spent doing actual work. Ie, speculation, branch prediction, doing the actual ALU math, cache management, prefetching, etc... Almost none of it goes to dealing with the ISA. The almost singular advantage ARM has is just the looser memory model, although the best ARM CPU on the market by a landslide (Apple's M1) can run x86's memory model, and it seemingly doesn't cost much (hence how Rosetta 2 can perform so well)
- imtringued 5y agoModern ARM CPUs go so far as having entire systems that translate external ARM instructions to an internal risc based one. I hope you realize that any big instruction set can be reduced to a smaller instruction set. The only instruction set that cannot be reduced further is a single instruction set.