6 ms·
In some cases RISC-V ISA spec is definitely the one to blame: 1) https://github.com/llvm/llvm-project/issues/150263 https://github.com/llvm/llvm-project/issues
by newpavlov 7mo ago
In some cases RISC-V ISA spec is definitely the one to blame:
1) https://github.com/llvm/llvm-project/issues/150263 https://github.com/llvm/llvm-project/issues/150263
2) https://github.com/llvm/llvm-project/issues/141488 https://github.com/llvm/llvm-project/issues/141488
Another example is hard-coded 4 KiB page size which effectively kneecaps ISA when compared against ARM.
- adastra22 7mo agoAlso the bit manipulation extension wasn't part of the core. So things like bit rotation is slow for no good reason, if you want portable code. Why? Who knows.
- fidotron 7mo agoThe fact the Hazard3 designer ended up creating an extension to resolve related oddities was kind of astonishing. Why did it fall to them to do it? Impressive that he did, but it shouldn't have been necessary.
- rllj 7mo agoWhich extension is that?
- mjmas 7mo agoAn extension he calls Xh3bextm. For extracting multiple bits from bitfields. https://wren.wtf/hazard3/doc/#extension-xh3bextm-section https://wren.wtf/hazard3/doc/#extension-xh3bextm-section There are also four other custom extensions implemented.
- wren6991 7mo agoThis extension wasn't strictly necessary but it makes decode of Arm instructions faster in the bootrom's Arm emulator.
- adgjlsfhk1 7mo ago> Also the bit manipulation extension wasn't part of the core. This is primarily because core is primarily a teaching ISA. One of the best parts about RiscV is that you can teach a freshman level architecture class or a senior level chip building project with an ISA that is actually used. Anything powerful to run (a non built from source manually) linux will support a profile that bundles all the commonly needed instructions to be fast.
- hackyhacky 7mo ago> One of the best parts about RiscV is that you can teach a freshman level architecture class or a senior level chip building project with an ISA that is actually used. Same could be said of MIPS. My understanding is the RISC-V raison d'etre is rather avoidance of patented/copywritten designs.
- adgjlsfhk1 7mo agothe avoidance of patent/copyright is critical for (legally) having students design their own chips. MIPS was pretty good (and widely used) for teaching assembly, but pretty bad for teaching a class where students design chips
- userbinator 7mo agoMIPS patents have long expired too (and incidentally for any other CPU released prior to 2006), so that's a moot point.
- musicale 7mo agoThis is largely contradicted by the (pre RISC-V) MIPS editions of Patterson & Hennessy, Harris & Harris, etc., which teach you how to design a MIPS datapath (at the gate level.) Regarding silicon implementations, consider that 1) you can synthesize it from HDL/RTL designs using modern CAD tools, and 2) MIPS was originally designed to be simple enough for grad students to implement with the primitive CAD tools of the 1980s (basically semi-manual layout).
- 7mo ago
- mort96 7mo agoDo you typically care about portability to the degree that you want the same machine code to execute on both a Linux box and a microcontroller? Why?
- weebull 7mo agoAll of those things are solved with modern extensions. It's like comparing pre-MMX x86 code with modern x86. Misaligned loads and stores are Zicclsm, bit manipulation is Zb[abcs], atomic memory operations are made mandatory in Ziccamoa. All of these extensions are mandatory in the RVA22 and RVA23 profiles and so will be implemented on any up to date RISC-V core. It's definitely worth setting your compiler target appropriately before making comparisons.
- edflsafoiewq 7mo agoWhat about page size?
- ori_b 7mo agoIt's 4k on x86 as well. Doesn't seem to hurt so bad -- at least, not enough to explain the risc-v performance gap.
- twoodfin 7mo agoHmm? x86 has supported much larger “huge” page sizes for ages.
- ori_b 7mo agoYes, and Linux. at least historically, has not used them without explicit program opt-in. Often advice is to disable transparent huge pages for performance reasons. Not sure about other operating systems. See, for example, https://www.pingcap.com/blog/transparent-huge-pages-why-we-disable-it-for-databases/ https://www.pingcap.com/blog/transparent-huge-pages-why-we-d...
- jorvi 7mo agoHuh, no? The usual advice is to enable THPs for performance, you only disable them in specific scenarios.
- 7mo ago
- direwolf20 7mo agoThe first one is common across many architectures, including ARM, and the second is just LLVM developers not understanding how cmpxchg works
- tosti 7mo agoRegarding misaligned reads, IIRC only x86 hides non-aligned memory access. It's still slower than aligned reads. Other processors just fault, so it would make sense to do the same on riscv. The problem is decades of software being written on a chip that from the outside appears not to care.
- pjmlp 7mo agoOn modern CPUs, it used not to be something to care about in the past across 8, 16, 32 bit generations, outside RISC.
- inkyoto 7mo agoPDP-11, m68k – to name a few, did not allow misaligned access to anything that was not a byte. Neither are RISC nor modern.
- pjmlp 7mo agoIn regards to 68000 I don't remember, only used it during demoscene coding parties when allowed to touch Amiga from my friends. I have only seen PDP-11 Assembly snippets in UNIX related books, wasn't aware of its alignment requirements.
- inkyoto 7mo agoPDP-11 was a major source of inspiration for m68k architecture designers. The influence can be seen in multiple places, starting from the orthogonal ISA design down to instruction mnemonics. It is quite likely that not allowing the misaligned access was also influenced by PDP-11.
- fredoralive 7mo agoARM Cortex-A cores also allow unaligned access (MCU cores don't though, and older ARM is weird). There's perhaps a hint if the two most popular CPU architectures have ended up in the forgiving approach to unaligned access, rather than the penalising approach of raising an interrupt.
- torginus 7mo agoUnaligned load/store is a horrible feature to implement. Page size can be easily extended down the line without breaking changes.
- GoblinSlayer 7mo ago> 1) https://github.com/llvm/llvm-project/issues/150263 https://github.com/llvm/llvm-project/issues/150263 Huh? They have no idea what they are doing. If data is unaligned, the solution is memcpy, not compiler optimizations, also their hack of 17 loads is buffer overflow. Also not ISA spec problem.