4 ms·
Sad that there was no mention / evaluation of RISC-V in this post, which attempts to resolve exactly the problems he identifies...
by maaku 10y ago
Sad that there was no mention / evaluation of RISC-V in this post, which attempts to resolve exactly the problems he identifies...
- microcolonel 10y agoI don't think RISC-V really addresses the encoding inefficiency problem, except for the "C" extension, sorta. Though I don't think that for OoO superscalar architectures, icache pressure is as much of a problem as it is on a fancy vliw. But yeah, would be nice to get a take on RISC-V in context of this rant.
- _chris_ 10y agoHuh? RISC-V with the compressed extension is incredibly efficient in its encoding. Better than x86 or ARM in both static and dynamic bytes per program. Also, Icache pressure is a huge problem in modern warehouse-scale computers. Any processor that cares about performance will almost certainly be implementing the C extension to RISC-V. It also enables more efficient macro-op fusion, turning common two instruction 4-byte idioms into a single, more powerful instruction.
- microcolonel 10y agoThanks for going into more detail. I was basing my assumption that it wasn't a huge problem on the fact that the only people who complain about it first seem to the folks designing the Mill. They have a ridiculous/insane/cool solution to it. Everyone else seems to first mention their cool branch predictor, or vector processor.
- renox 10y agoGiven the lack of CCR in the RISC-V I doubt that he would be very impressed by it's ease of use..
- dietrichepp 10y agoI'm not convinced that's a big deal. You just end up using a register of your choice and sticking a flag in it.
- renox 10y agoOK, please show me the code to do a long addition or a long multiplication in RISC-V. (long as in 'multiple words')
- dietrichepp 10y agoHere. 64-bit addition on RV32I. ; input 1 (msb r1, lsb r2) ; input 2 (r3, r4) ; output (r5, r6) xori r5, r4, -1 sltu r5, r5, r2 add r6, r4, r2 add r5, r5, r3 add r5, r5, r1 This is what I mean. Outside a few applications (mostly asymmetric crypto) nobody cares that it takes five instructions instead of two. Remember that this is the same processor that outright omits multiplication from the core spec.