5 ms·
> This pursuit of minimalism has resulted in false orthogonalities (such as reusing the same instruction for branches, calls and returns) They are all kinds of
by outsomnia 6y ago
> This pursuit of minimalism has resulted in false orthogonalities (such as reusing the same instruction for branches, calls and returns)
They are all kinds of changes in flow, it's not crazy to have an instruction for that primitive.
He complains Risc-V need 4 instructions to do what x86_64 and arm does in two, but... it says Risc-V. And x86_64 CISC instructions devolve to a pile of microcode anyway.
Guy invested in powerful incumbant using a completely different and even more established enumbant to bash the challenger over the head with doesn't feel like I am learning anything useful when I read it.
- nojs 6y agoThe author is a woman fyi :)
- K2L8M11N2 6y agoAccording to their Twitter, the author uses they/them pronouns. https://twitter.com/erincandescent https://twitter.com/erincandescent
- Taniwha 6y agoThey got the riscv code slightly wrong (but the instruction count correct). (puts on chip designer's hat) Essentially the ARM/x64 case turns the load memory address calculation into a 4-input adder (so 2 layers of adders) and maybe a couple of extra gates because there are multiple addressing modes. Riscv's equivalent is a 2-input adder. Those get into a critical cache (and TLB) access path and that limits how fast your CPU's core clock can be (or forces you to split that path into 2 clocks). Essentially that's part of the whole RISC idea - simple means faster - you can run your core clocks faster if the decode (and address calculations etc etc) are simpler - getting rid of lots of addressing modes was a big part of the original RISC movement. I think all 4 of those riscv instructions are 16-bit ones so they may even fit into the same space as the 2 ARM ones (haven't hacked on ARM for a while) BTW chances are that that x86_64 mov instruction is not being devolved into more than one internal uOp (might be two if they separate off the address calculation into it's own uOp)
- Veedrac 6y ago> Essentially that's part of the whole RISC idea - simple means faster - you can run your core clocks faster if the decode (and address calculations etc etc) are simpler - getting rid of lots of addressing modes was a big part of the original RISC movement. Nobody believes this anymore, not even the RISC guys. Look at Apple chips; the advantage of a simpler instruction set is width (degree and depth of superscalar execution) and the ability to put more optimizations in hardware. Clocks are all bound by roughly the same limits nowadays.
- Taniwha 6y agoI'd disagree - but I'm talking about the difference between a 1Ghz clock and a 5Ghz clock - it's always going to matter at the cutting edge - if the clocks are equal then what you're trading off is pipeline depth - which effects all sorts of stuff like the cost of mispredicted branches (and as a result the sizes and complexity of branch predictors)
- Tuna-Fish 6y ago> Essentially that's part of the whole RISC idea - simple means faster The problem with this is that today, your clock speed is bound by neither your decode nor your ALUs. Only implementing weaker ALUs makes sense from an optimization standpoint if it buys you more clock speed. But as it doesn't, it just leaves you competing with another CPU that has the same clock speed as you do, and which does a lot more per clock than you do.
- ip26 6y agoTurns out faster clocks don't help much if your logic isn't getting faster too. You do less work per cycle, while your cycle overhead remains fixed, resulting in less work done per unit time. If it wasn't for branch misprediction, you'd be better off with very deep pipes and slower clocks - see GPUs.
- ithkuil 6y agoI think the point of GP is that not every instruction makes use of the more complicated (4 input) adder, so not every instruction should have to pay for the latency cost associated. But at other pointed out this only makes sense if the ALU is the critical path and that having a two level adder significantly impacts the latency
- PeCaN 6y ago>[She] complains Risc-V need 4 instructions to do what x86_64 and arm does in two, but... it says Risc-V. So… what, it should take 5 instructions? Executing more instructions for a (really) common operation doesn't mean an ISA is somehow better designed or "more RISC", it means it executes more instructions. >And x86_64 CISC instructions devolve to a pile of microcode anyway. Some people seem to have this impression that like every x86 instruction is implemented in microcode (very, very few of them are) and even charitably interpreting that as "decodes to multiple uops" (which is completely different) is still not right. The mov in the example is 1 uop.
- tom_mellior 6y ago> Executing more instructions for a (really) common operation doesn't mean an ISA is somehow better designed or "more RISC", it means it executes more instructions. True. But as bonzini points out (or rather, hints at) in https://news.ycombinator.com/item?id=24958644 https://news.ycombinator.com/item?id=24958644, the really common operation for array indexing is inside a counted loop, and there the compiler will optimize the address computation and not shift-and-add on every iteration. See https://gcc.godbolt.org/z/x5Mr66 https://gcc.godbolt.org/z/x5Mr66 for an example: for (int i = 0; i < n; i++) { sum += p[i]; } compiles to a four-instruction loop on x86-64 (if you convince GCC not to unroll the loop): .L3: addsd xmm0, QWORD PTR [rdi] add rdi, 8 cmp rax, rdi jne .L3 and also to a four-instruction loop on RISC-V: .L3: fld fa5,0(a0) addi a0,a0,8 fadd.d fa0,fa0,fa5 bne a5,a0,.L3 This isn't a complete refutation of the author's point, but it does mitigate the impact somewhat.
- PeCaN 6y agoThat's fair. It's definitely not a killer, (or even in my opinion the worst thing about RISC-V,) just another one of these random little annoyances that I'm not really sure why RISC-V doesn't include.
- ncmncm 6y agoOne common use of array indexing walks the array sequentially. But hash tables are used here and there, also in loops. Some people know them as "dictionaries" or "key/value stores".
- Jasper_ 6y agoFrom a software perspective, changes in flow are identical. But from a hardware standpoint, local jumps, indirect calls and returns are all predicted differently. The spec actually has suggested forms for each kind (the author points them out in the article), so a good portion of encoding space is used on a large number of variants that will have poor performance and never be seen in practice. Especially when those bits could be used for far better purposes.
- admax88q 6y agoHaving more instructions for common tasks also puts more pressure on your memory bandwidth and instruction cache. Execution time alone is not the only factor in performance.