4 ms·
One of the factors is that all ARM64 instructions are the same length which makes decoding simpler.
by 58028641 4y ago
One of the factors is that all ARM64 instructions are the same length which makes decoding simpler.
- mnw21cam 4y agoThis is the correct answer. RISC has had its moments. Way back, it was better than CISC because the simpler instructions allowed a higher clock rate. Then CISC CPUs turned into RISC CPUs with a CISC-to-RISC translation on the front end. With that in place, there was a whole stretch of time where the only real advantage that RISC had over CISC was that RISC didn't have to have a lump of silicon that translates CISC into RISC instructions. However, now there's enough space on the die for a CPU to have lots of parallel execution units, and the part of it that became really difficult to scale was the CISC-to-RISC translation unit, because each CISC instruction had an unpredictable length, making working out which instructions you can translate tricky and silicon-consuming. And so RISC has a significant advantage once more, and that is the ability to vastly-simplify the part of the CPU that feeds instructions into the execution pool, compared to a CISC CPU, because the instructions are fixed-length. This allows this part to translate more instructions per clock cycle than a typical CISC CPU can, and this is what gives the improved performance.
- nomel 4y agoI'm a dummy. Do I have this right? > CISC-to-RISC translation on the front end I would naively assume that this would be an advantage, since you could easily change the hardware used for any CISC instruction, finding better ways to make it faster. The "work unit" is more abstract, so you could throw the whole problem at dedicated silicon. Or, you could remove dedicated silicon, and just have the CIST spit out a list of RISC instructions. It seems that, for RISC, you could never throw a more abstract "work unit" at dedicated silicon, without buffering instructions, to see if the intent matches the accelerators. Chip specific compilers would almost be required, to handle the abstraction.
- mnw21cam 4y agoIt was an advantage. Now it's a liability, because it is hard to scale the CISC-to-RISC translation up to decode many instructions per clock cycle. The translation unit has become the bottleneck.
- nomel 4y agoWhy? Wouldn’t it be adding a list of instructions to the queue, for the RISC side? Why would that RISC side have to be slower? I would assume it would be mostly independent. Or is that queue pollution the problem, rather than the execution of it?
- mnw21cam 4y agoThe problem is that CISC instructions are variable-length, so it's easy to work out how to decode the first instruction, but the second instruction depends on the length of the first instruction, and if you try to decode four instructions at once in a single clock cycle then it all gets a bit too much, which reduces your maximum clock speed. In comparison, a RISC instruction decoder knows that each instruction is the same length, so each instruction can be decoded without depending on the ones before it. This simplifies the decoder so much that it makes it possible to decode four instructions per clock cycle without investing in too much silicon to do it, and while keeping a high clock speed.
- 3pm 4y agoCISC-to-RISC overhead is not the only factor probably. Snapdragons for example don't have CISC-to-RISC translation, yet they seem to under perform both Intel and M1.
- mnw21cam 4y agoSnapdragons aren't aiming for 4 instructions decoded per clock cycle.
- 3pm 4y agoI was trying to say that if CISC-to-RISK overhead is the main contributing factor then other laptops without the overhead would be competitive. ThinkPad X13s with a 4 Core Snapdragon for example. It is a pure-RISK machine and it seem to under perform Intel and is almost twice as slow as M1.
- mnw21cam 4y agoRISC is the secret sauce that allows M2 to decode four instructions per clock cycle, due to the constant instruction length, and that is what gives the M2 a speed advantage over CISC instruction sets. The snapdragon CPUs aren't trying to decode four instructions per clock cycle, so they aren't taking advantage of that feature of RISC instructions.
- 3pm 4y agoGot it. It is not enough to have a fixed width instruction set, you also have to actually decode and execute them faster after decoding. Wonder at what point we will start seeing competitive ARM CPUs.