4 ms·
I largely agree with you, but funnily enough the very same blog has a great post on the x86 decoding myth https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-d
by zorgmonkey 2y ago
I largely agree with you, but funnily enough the very same blog has a great post on the x86 decoding myth https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-matter/ https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
- trueismywork 2y agoI'm not sure I agree with that. The unknown length of instructions in x86 does make decoders more power hungry. There's no doubt about that. That is a big problem for power efificency of x86 and the blog doesn't address that at all. M1 really is a counterexample to all theory that Jim is saying etc. The real proof would be if same results were also reproduced on M1 instead of Zen
- oelang 2y agoJim was involved in the early versions of Zen & M1, I believe he knows. Apples M series looks very impressive because typically, at launch, they are node ahead of the competition, so early access deals with TSMC is the secret weapon this buys them about 6 months. They also are primarily laptop chips, AMD has competitive technology but always launches the low power chips after the desktop & server parts.
- Jensson 2y ago> so early access deals with TSMC is the secret weapon this buys them about 6 months Aren't Apple typically 2 years ahead? M1 came out 2020, other CPUs from the same node level (5 nm TSMC) came out 2022. If you mean apple launches their 6 months ahead of the rest of the industry gets on the previous node, sure, but not the current node. What you are thinking about is maybe that AMD 7nm is comparable to Apple 5nm, but really what you should compare is todays AMD cpus with the Apple cpu from 2022, since they are on the same architecture. But yeah, all the impressive bits about Apple performance disapears once you take architecture into account.
- torginus 2y agoNot necessarily. Qualcomm just released its Windows chips, and in the benchmarks I've seen, it loses to the M1 in power efficiency, despite being built on a more advanced node, performing much closer the the Intel and AMD offerings. Apple is just that good.
- deleted 2y ago[deleted]
- tremon 2y agoThe unknown length of instructions in x86 does make decoders more power hungry. There's no doubt about that. I have doubts about that. I-cache word lines are much larger than instructions anyway, and it was the reduction in memory fetch operations that made THUMB more energy-efficient on ARM (and even there, there's lots of discussion on whether that claim holds up). And if you're going for fixed-width instructions then many instructions will use more space than they use now, reducing the overall effectiveness of the I-cache. So even if you can prove that a fixed-size decoder uses less power, you will still need to prove that that gain in decoder power efficiency is greater than the increased power usage due to reduced instruction density and accompanying loss in cache efficiency.
- weebull 2y agoIt's the width of multiplexing that has to on between having a fetch line and extracting a single instruction. As an instruction can start at many different locations, you need to be able to shift down from all those positions. That's not too bad for the first instruction in a line but the second instruction is dependant on how the first instruction decides, and the third dependent on the second. Etc. So it's not only a big multiplexer tree, but a content dependent multiplexer tree. If you're trying to unpack multiple instructions each clock (or course you are. You've got six schedulers to feed) then that's a big pile of logic. Even RISC-V has this problem, but there they've limited it to two sizes of instruction (2 and 4 bytes), and the size is in the first 2 bits of each instruction (so no fancy decode needed)