4 ms·
And M1 does the same. All modern CPU's are really similar internally in that sense. ISA is just a frontend that gets translated into micro-ops that then get sch
by sharpneli 6y ago
And M1 does the same. All modern CPU's are really similar internally in that sense. ISA is just a frontend that gets translated into micro-ops that then get scheduled based on dependencies and available execution ports. Even the registers in ISA don't match the internal registers at all. ARM64 has 32 general purpose registers in ISA level. M1 seems to have 354 internally [Anandtech]
This is also the reason why the whole CISC and RISC debate in it's original form is outdated. The processors internally are all RISC. But the ISA can be more complex.
The x86 ISA makes the decoder harder to parallelize, so it takes more chip area compared to equivalent width for ARM64. And the wider you want to go the harder it becomes, whereas with ARM64 you just slap more decoders.
Another is the x86 memory model that restricts how stores can be issued into memory so that they're visible to other cores.
This is also a good thing for AMD. They could "just" make a Zen ARM CPU. Sure it would be a lot of work, but vast majority of the chip is shared.
- hajile 6y agoThat frontend isn't free and it isn't small. Look at modern Intel processors and you'll see that the decoder takes as much die area as the entire Integer ALU (if you don't count caches). Unlike the ALU which powergates unused ports, the decoder almost never turns off. The more 1-to-1 your translations to uops are, the less power and die area you need to spend decoding them. In addition, less complex translations means fewer pipeline stages needed for the same design which also has lots of ramifications.
- sharpneli 6y agoYap. And ARM has the advantage of requiring a smaller frontend. Especially when one looks at wider decoders. On the other hand if your ISA is the micro-ops directly then the instructions start to take ridiculous amounts of space. It's a balancing act between instruction size (to save instruction cache) and the complexity of decoding them. And it's not just about being 1:1. It's also about how wide you can go. And variable length encodings simply are fundamentally more hard to parse in parallel fashion. That means a wider unit is harder to achieve, needing more space and power.