5 ms·
It's true that ARM64 has a load-store architecture and fixed-length instructions (the latter depending on the former for encoding space efficiency). Other than
by psykotic 6y ago
It's true that ARM64 has a load-store architecture and fixed-length instructions (the latter depending on the former for encoding space efficiency). Other than that, the instruction set design is very far from minimalist textbook-style RISC ISAs like RISC-V. It has both flag-based branches and fused compare-to-zero-and-branch instructions. It has very complex immediate encodings. It has instructions for loading/storing register pairs. It has pre-increment/post-increment addressing modes of the kind that were hallmarks of CISCs like M68K and VAX.
It seems unwise to draw far-reaching conclusions about RISC-V or even ARM64's intrinsic merits versus Apple's CPU designers when there are so many variables. The frontend decoder hasn't been a frequent bottleneck in Intel cores for a long time and they could scale it up more aggressively if they wanted.
Apple's engineers did a great job. That seems to be the conclusion we can draw based on currently available evidence.
- IshKebab 6y agoI'm not sure what your point is, but no modern ISA is really bare-bones RISC. They're all somewhere in the middle, including RISC-V, despite the name (it just puts the more complex instructions in optional extensions).
- FullyFunctional 6y ago> The frontend decoder hasn't been a frequent bottleneck in Intel cores for a long time and they could scale it up more aggressively if they wanted. This isn't grounded in any facts. Decoding the variable length x86 ISA costs you exponentially in decoding width, both power and area. You can scale it, but it will never be efficient. The way Intel and AMD combat this is by having a decoded uOP cache from which the issue width is typically twice that of the frontend decoder. Arm64 has an inherent advantage here (RISC-V does not have quite the same advantage as RV64GC instructions are a mix of 16- and 32-bit). Arm64 also is much more recent design than x86_64 that learned a lot from the past experience and isn't bogged down by a lot of useless legacy. This helps. Arm64 is rather large for a RISC ISA, but it's mostly pretty good (however IMO RISC-V's lack of flags and implementation of conditional branches is superior).
- psykotic 6y agoOf course a fixed-length ISA has an inherent advantage for parallel decoding efficiency. The question is whether that is a decisive advantage in M1's impressive performance. After Intel refined their decoder and uop cache, you virtually never see that part of the frontend as a bottleneck when doing microarchitectural profiling. That's been true since Sandy Bridge but even more so since Skylake. All the legacy junk in x86 is obviously a pain for Intel to support. Any blank-slate ISA is going to have an advantage there.
- amelius 6y agoIn any case, it would be relatively simple for intel/amd engineers to evaluate the effect of different parameters using their quantitative analysis tools which include an emulation environment. I don't think it makes much sense to speculate here about these parameters.
- twic 6y agoPresumably the importance of decode bandwidth depends on what you're decoding. Most classic computationally intensive work (video encoding, science, but also benchmarks) spends its time in fairly tight loops or small kernels, running over large data. uop caches make decode bandwidth irrelevant here. But general usage of a machine sees the instruction pointer wander all over the place (particularly if you have multiple tabs of JavaScript open). More decode bandwidth means more performance here. Are compilers are an an example of a heavy workload with a large hot code size? It would be interesting to compare the M1's advantage in compiling to its advantage in, say, video encoding.
- Symmetry 6y agoIt doesn't take exponential power. My understanding is that the basic approach for instructions without boundary tagging in L1I$ is to start decoding every byte in the stream in parallel, discard the ones that don't make sense, and then later propagate boundary to boundary across the length of the fetch window. Sort of similar to how a carry-bypass adder works. This is expensive but not that expensive compared to other structures. But it does mean that x86 designs tend to carefully balance the size of the decoders to other structures to make sure they're not the binding constraint too often. With ARM the approach seems to be more to make the front end 50% bigger than you think you need to be sure it's never a problem and refill the front end buffers more quickly after a mispredict.
- Symmetry 6y agoThe traditional RISC philosophy stemmed from the constraints on chip development that mostly existed from the 80s to the mid 90s. After that ballooning transistor counts and design effort for top line out of order application processors made reducing the number of instructions pretty pointless in that design space, though limiting the number of ways instructions could interact through a load-store architecture and keeping decode simple through fixed length instructions (plus longer jumps) remain relevant useful things to take away from RISC. All the complexities that ARM has let it do more with fewer instructions and a high performance RISC-V core is going to have to do enough instruction fusion that its internal ops will end up being just as complicates as those that ARM uses, but it'll also have the disadvantage of having to do that extra fusion. But of course if the target isn't a high end application processor but instead a microcontroller, say, RISC-V's simplicity has a lot going for it. Or for a grad student implementing a simple OoO processor in a semester long class. Or back when I was doing my thesis having an open source core to modify would have been a huge advantage. As the article says RISC-V can be a success without replacing ARM, POWER, and x86.