4 ms·
Apples chips utterly dominate in performance per watt versus x86. Arm64s simple fixed length instruction set is absolutely central in this. Qualcomm now exceedi
by fooblaster 2y ago
Apples chips utterly dominate in performance per watt versus x86. Arm64s simple fixed length instruction set is absolutely central in this. Qualcomm now exceeding Intel in their mobile chips is further evidence ISA is critically important. In the data center this is also happening with graviton 4 and eventually grace hopper.
Anyway, my point isn't that implementation does not effect performance. Of course it does. My point was that Linus has asserted ISA is irrelevant to performance, and that only implementation matters. That assertion does not hold up.
- snvzz 2y ago>Arm64s simple fixed length instruction set is absolutely central in this. Note that doing variable instruction length is not an issue per-se as proven by more dense, yet still easy to decode RISC-V. The way x86 does it, however, is very costly to decode and overall awful.
- fooblaster 2y agoyeah, there's a lot of nuance there. Not all variable length isas are a problem. X86 definitely is.
- brucehoult 2y agoTLDR: RISC-V 2-byte and 4-byte dual instruction lengths have no significant decoding scaling / latency problem at any practical CPU width. Two instruction widths gives big benefits in reduced code size, better cache utilisation, less instruction fetch bandwidth etc and at very low cost. RISC-V instruction length decoding requires you to look at only 2 bits per 2 bytes of code (an AND gate on those two bits reduces them to 1 bit). So then decoding, say, 32 bytes of code -- which will contain between 8 and 16 instructions with an average around 11 -- requires looking at only 16 signals. That's close to lookup table territory. In fact the way to do it (at least what I came up with) is to make a decoder for each group of 4 aligned bytes, plus optionally the previous 2 bytes (so 6 bytes in total). Each block can decode either one instruction (an aligned 4-byte instruction, or an unaligned 4 byte instruction starting in the previous block, followed by a second unaligned 4-byte instruction which is left for the next decode) or two instructions (two 2-byte instructions, or an unaligned 4-byte instruction started in the previous block, followed by a 2-byte instruction). Each block receives a signal saying whether or not the previous block consumed the last two bytes, and outputs a similar signal to the next decoder block. So you can daisy-chain these, with 8 such decoder blocks in parsing 32 bytes of code. That's not a HUGE delay. But you can do better. Each block generates two signals, one saying whether the last two bytes of the group are consumed if the previous decoder block consumes all of its bytes, and a second signal saying whether the last two bytes of the group are consumer if the previous decoder block does NOT consume its last two bytes of code. These signals can be thought of as "generates" and "propagates" signals, analogous to a full-adder's carry output. These two signals can be generated in parallel, and INDEPENDENT OF all previous blocks. You can even decode the one or two complete instructions both possible ways, using three decoders: a 4-byte only decoder started at -2 bytes, a 2-byte only decoder starting at +2 bytes, and a decoder that can do both 2-byte and 4-byte instructions starting at +0. ALL INDEPENDENT of any previous instruction decoder blocks. And then you can use the same technique as in a carry-lookahead adder to select which decoded output (one or two instructions) you use from each decoder block, with the same scaling properties as in an adder. Except that if you're decoding only 32 bytes of code per cycle then you're looking at the analogue of an 8 bit adder, where carry-lookahead is barely worth it. The technique comes into its own in 16, 32, 64, 128 bit adders. That's equivalent to decoding 64, 128, 256, 512 bytes of code PER CYCLE, which is far beyond the bounds of a practical CPU because of basic block lengths and predicting multiple branches. The cost is that each 4 bytes of code needs three decoders while producing maximum two instructions, though the 3rd decoder only has to deal with the limited 2-byte C extension instruction set (and only one of the other two decoders needs a C decoder). Note: you can save on decoders, at the expense of a little latency, by having only one full-strength (2-byte and 4-byte) decoder and one 2-byte only decoder, and using the "carry" input to MUX either bytes -2..1 or bytes 0..3 into the full-strength decoder. Yes, it's more complex than an Aarch64 decoder of the same width, but it's completely manageable and massively simpler than x86.
- fooblaster 2y agoyeah, it's definitely manageable when the instructions don't vary in size from 1-15 bytes. interesting details.