7 ms·
All of those things are solved with modern extensions. It's like comparing pre-MMX x86 code with modern x86. Misaligned loads and stores are Zicclsm, bit manipu
by weebull 7mo ago
All of those things are solved with modern extensions. It's like comparing pre-MMX x86 code with modern x86. Misaligned loads and stores are Zicclsm, bit manipulation is Zb[abcs], atomic memory operations are made mandatory in Ziccamoa.
All of these extensions are mandatory in the RVA22 and RVA23 profiles and so will be implemented on any up to date RISC-V core. It's definitely worth setting your compiler target appropriately before making comparisons.
- edflsafoiewq 7mo agoWhat about page size?
- ori_b 7mo agoIt's 4k on x86 as well. Doesn't seem to hurt so bad -- at least, not enough to explain the risc-v performance gap.
- twoodfin 7mo agoHmm? x86 has supported much larger “huge” page sizes for ages.
- ori_b 7mo agoYes, and Linux. at least historically, has not used them without explicit program opt-in. Often advice is to disable transparent huge pages for performance reasons. Not sure about other operating systems. See, for example, https://www.pingcap.com/blog/transparent-huge-pages-why-we-disable-it-for-databases/ https://www.pingcap.com/blog/transparent-huge-pages-why-we-d...
- jorvi 7mo agoHuh, no? The usual advice is to enable THPs for performance, you only disable them in specific scenarios.
- wren6991 7mo agoYep, RISC-V also has these megapages. 4k is the last-level page size. You get larger pages (4M on 32-bit and 2M/1G on 64-bit) by terminating the walk at higher levels of the page table.
- jabl 7mo agox86 has decades of knowhow and a zillion transistors to spend on making the memory pipeline, TLB caching & prefetching etc. etc. really really good. They work as well as they do despite the 4k base page size, not because of it. If you'd start from a clean sheet today you'd probably end up with a somewhat bigger base page size. Not hugely larger though, as that wastes a lot of memory for most applications. Maybe 16k like some ARM chips use?
- rwmj 7mo agoRISC-V has the Svnapot extension for large page sizes https://riscv.github.io/riscv-unified-db/manual/html/isa/isa_20240411/exts/Svnapot.html https://riscv.github.io/riscv-unified-db/manual/html/isa/isa...
- LeFantome 7mo agoUbuntu being RVA23 is looking smarter and smarter. The RISC-V ecosystem being handicapped by backwards compatibility does not make sense at this point. Every new RISC-V board is going to be RVA23 capable. Now is the time to draw a line in the sand.
- saagarjha 7mo agoI’d be kind of depressed if every new RISC-V board was not RVA23 capable.
- newpavlov 7mo ago>Misaligned loads and stores are Zicclsm Nope. See https://github.com/llvm/llvm-project/issues/110454 https://github.com/llvm/llvm-project/issues/110454 which was linked in the first issue. The spec authors have managed to made a mess even here. Now they want to introduce yet another (sic!) extension Oilsm... It maaaaaay become part of RVA30, so in the best case scenario it will be decades before we will be able to rely on it widely (especially considering that RVA23 is likely to become heavily entrenched as "the default"). IMO the spec authors should've mandated that the base load/store instructions work only with aligned pointers and introduced misaligned instructions in a separate early extension. (After all, passing a misaligned pointer where your code does not expect it is a correctness issue.) But I would've been fine as well if they mandated that misaligned pointers should be always accepted. Instead we have to deal the terrible middle ground. >atomic memory operations are made mandatory in Ziccamoa In other words, forget about potential performance advantages of load-link/store-conditional instructions. `compare_exchange` and `compare_exchange_weak` will always compile into the same instructions. And I guess you are fine with the page size part. I know there are huge-page-like proposals, but they do not resolve the fundamental issue. I have other minor performance-related nits such `seed` CSR being allowed to produce poor quality entropy which means that we have bring a whole CSPRNG if we want to generate a cryptographic key or nonce on a low-powered micro-controller. By no means I consider myself a RISC-V expert, if anything my familiarity with the ISA as a systems language programmer is quite shallow, but the number of accumulated disappointments even from such shallow familiarity has cooled my enthusiasm for RISC-V quite significantly.
- IshKebab 7mo agoI think having separate unaligned load/store instructions would be a much worse design, not least because they use a lot of the opcode space. I don't understand why you don't just have an option to not generate misaligned loads for people that happen to be running on CPUs where it's really slow. You don't need to wait for a profile for that. As for `seed`, if you're running on a microcontroller you can just look up the data sheet to see if it's seed entropy is sufficient. By the time you get to CPUs where portable code is important a CSPRNG is probably fine. I agree about page size though. Svnapot seems overly complicated and gives only a fraction of the advantages of actually bigger pages.
- deleted 7mo ago[deleted]
- sidewndr46 7mo agoYou're correct but I guess my thoughts are if we're going to wind up with a mess of extensions, why not just use x86-64?
- whaleofatw2022 7mo agoBecause the ISA is not encumbered the way other ISAs are legally, and there are use cases where the minimal profile is fine for the sake of embedded whatever vs the cost to implement the extensions
- computably 7mo ago> why not just use x86-64? Uh, because you can't? It's not open in any meaningful sense.
- userbinator 7mo agoThe original amd64 came out in 2003. Any patents on the original instruction set have long expired, and even more so for 32-bit x86.
- panick21_ 7mo agoIts not about patents. Believe what you want but there is a reason nobody else is doing x86 or ARM chips unless they are allowed by the owner.
- dbdr 7mo agoYou're probably right. It would be helpful to say what the reason is, if it's not patents.
- panick21_ 7mo agoI'm not a lawyer but I would assume its copyright. Kind of like API in software. In software somehow this does not apply most of the time. But it seems in hardware this is very real. But I would appreciate a lawyer jumping in. I know for example that Berkley when thinking pre-RISC-V that they had a deal with Intel about using x86-64 for research. But they were not able to share the designs.
- cmovq 7mo agoBut RISC-V is a _new_ ISA. Why did we start out with the wrong design that now needs a bunch of extensions? RISC-V should have taken the learnings from x86 and ARM but instead they seem to be committing the same mistakes.
- hun3 7mo agoIt was kind of an experiment from start. Some ideas turned out to be good, so we keep them. Some ideas turned out not to be good, so we fix them with extensions.
- pjmlp 7mo agoThe problem with hardware expirements is that people owning the hardware are stuck with experiments.