2 ms·
> RISC-V is designed with great care taken to weight all decisions as not to hamper any scope of implementation, from the lowest power microcontrollers to the f
by brigade 4y ago
> RISC-V is designed with great care taken to weight all decisions as not to hamper any scope of implementation, from the lowest power microcontrollers to the fastest supercomputers, and everything in between.
RISC-V was consistently designed so that the simplest implementations remained as simple as possible. When design tradeoffs meant more complexity for more performant designs, any possible detriment was determined to be unimportant.
Notoriously, this is the whole "big cores can fuse instructions; that's got to be free for them, right?" Though the V extension also has a couple of fun ones with re-using v0 for masking (oh a big core can just keep track of whether it's a mask or not and switch which set of registers it's renaming from) and vsetvl (yeah let's make a big core speculate what effect it'll have on subsequent instructions)
> aarch64's awful code density
All the code density comparisons I've seen have been static. Do you know of any dynamic comparisons? I suspect 10% less dense code than RISC-V C (but still denser than x86-64) isn't the primary reason for having 3x more L1I than I think any RISC-V or x86 design currently has...
And even then, ARM for instance decided it was worth spending 50% of L1I cache area on a MOP cache for various reasons, a key one being that their implementation shaves off a cycle in branch mispredicts. If there's no predecode info stored in cache, I can easily see a predecode stage adding a cycle to the pipeline depth over assuming fixed-length. And if it is stored in SRAM you lose some of the codesize benefits...