16 ms·
RISC-V Is Sloooow
- rbanffy 7mo agoDon't blame the ISA - blame the silicon implementations AND the software with no architecture-specific optimisations. RISC-V will get there, eventually. I remember that ARM started as a speed demon with conscious power consumption, then was surpassed by x86s and PPCs on desktops and moved to embedded, where it shone by being very frugal with power, only to now be leaving the embedded space with implementations optimised for speed more than power.
- dmitrygr 7mo agoIF you care to read the article, they indeed do not blame the architecture but the available silicon implementations.
- rbanffy 7mo agoI did read it. A Banana Pi is not the fastest developer platform. The title is misleading. BTW, it's quite impressive how the s390x is so fast per core compared to the others. I mean, of course it's fast - we all knew that. And don't let IBM legal see this can be considered a published benchmark, because they are very shy about s390x performance numbers.
- menaerus 7mo agoWhich risc-v implementation is considered fast?
- patchnull 7mo ago[flagged]
- menaerus 7mo agoI remember taking down some notes wrt SiFive P870 specs, comparing them to x86_64, and reaching the same conclusion. Narrower core width (4-wide vs 8-wide), lower clock frequency (peaks at 3GHz) and no turbo (?), limited support for vector execution (128-bit vs 512-bit), limited L1 bandwidth (1x 128-bit load/cycle?), limited FP compute (2x 128-bit vs 2x 512-bit), load queue is also inconveniently small with 48 entries (affecting already limited load bandwidth), unclear system memory bandwidth and how it scales wrt the number of cores (L3 contention) although for the latter they seem to use what AMD is doing (exclusive L3 cache per chiplet).
- LeFantome 7mo agoSpacemiT K3 is about the same performance as a Rockchip RK3588. So, 4 years ago? Except the K3 kills it on AI (60 TOPS).
- NooneAtAll3 7mo agoDC-ROMA 2 is on the Rasperry 4 level of performance last I heard
- LeFantome 7mo ago> Which risc-v implementation is considered fast? SpacemiT K3 is 2010 Macbook performance single-core, 2019 Macbook Air multi-core, and better than M4 Apple Silicon for AI. So I guess it depends on what you are going to do with it.
- menaerus 7mo agoM4 is 38 TOPS at INT8 precision whereas SpacemiT K3 is 60 TOPS at INT4 precision so at best they would be equal in "AI" performance but they are not because the rest of the K3 chip is much less capable than M4 (as I would expect). E.g. M4 total system memory bandwidth is 120GB/s whereas K4 is 51GB/s, single core memory bandwidth is 100-120GB/s vs ~30GB/s. M4 has 10 CPU cores and neural engine with 16 cores whereas K3 has 8 CPU cores and 8 "AI" cores, K3 clock frequency is almost half the clock frequency in M4 etc. etc. But anyway thanks for sharing, always good to learn about new hardware.
- gt0 7mo agoI was really surprised by the s390x performance, but I also don't really understand why there are build time listed by architecture, not the actual processors.
- rbanffy 7mo agoProbably because that's just the infrastructure they have.
- pantalaimon 7mo agoi686 builds even faster
- kpil 7mo agoWhat's fast on Z platforms is typically IO rather than raw CPU - the platform can push a lot of parallell data. This is typically the bottleneck when compiling. The cores are in my experience moderately fast at most. Note that there are a lot of licencing options and I think some are speed-capped - but I don't think that applies to IFL - a standard CPU licence-restricted to only run linux.
- burntoutgray 7mo agoI thought I read somewhere that Z CPUs run at 5GHz ??
- Aurornis 7mo ago> A Banana Pi is not the fastest developer platform. What is the current fastest platform that isn’t exorbitantly expensive? Not upcoming releases, but something I can actually buy. I check in every 3-6 months but the situation hasn’t changed significantly yet.
- cestith 7mo agoWhat is the current fastest ppc64le implementation that isn’t exorbitantly expensive? How about the s390x?
- adgjlsfhk1 7mo agoA P550 based board is the best you can get for now (~2-3x faster than the Banana Pi). In 2-3 months there should be a number of SpaceMIT k3 chips that are ~4-6x faster than the banana pi and somewhat reasonably priced (~200-300). By the end of the year, however, you should be able to get an ascalon chip which should be way way faster than that (roughly apple m1/zen3 speed)
- snvzz 7mo ago>I did read it. A Banana Pi is not the fastest developer platform. The title is misleading. Ironically, its SoC (spacemiT K1) is slower than the JH7110 used in the first mass-produced RISC-V SBC, VisionFive 2. But unlike JH7110, it has vector 1.0, making it a very popular target. Of course, none of these pre-RVA23 boards will be relevant anymore, once the first development boards with RVA23-compatible K3 ship next month. These are also much faster than anything RISC-V currently purchasable. Developers have been playing with them for months through ssh access.
- tromp 7mo agoBut they didn't reflect that in a title like "current RISC-V silicon Is Sloooow" ...
- topspin 7mo agoI keep checking in on Tenstorrent every few months thinking Keller is going to rock our world... losing hope. At this point the most likely place for truly competitive RISC-V to appear is China.
- rbanffy 7mo ago> At this point the most likely place for fast RISC-V to appear is China. Or we just adopt Loongson.
- balou23 7mo agoTBH I still don't really get how it's different from MIPS. As far as I can tell... Loongson seems to be really just MIPS, while LoongArch is MIPS with some extra instructions.
- mananaysiempre 7mo agoBut legally distinct! I guess calling it M○PS was not enough for plausible deniability.
- genxy 7mo agoISAs shouldn't be patentable in the first place.
- pantalaimon 7mo agoThey did get rid of the delay slots and some other MIPS oddities
- bonzini 7mo agoLoongArch is, on a first approximation, an almost RISC-V user space instruction set together with MIPS-like privileged instructions and registers.
- mananaysiempre 7mo ago
- spiderice 7mo agoThen how do you justify the title?
- api 7mo agoA pattern I've noticed for a very long time: A lot of times the path to the highest performing CPU seems to be to optimize for power first, then speed, then repeat. That's because power and heat are a major design constraint that limits speed. I first noticed this way back with the Pentium 4 "Netburst" architecture vs. the smaller x86 cores that became the ancestor of the Core architecture. Intel eventually ran into a wall with P4 and then branched high performance cores off those lower-power ones and that's what gave us the venerable Core architecture that made Intel the dominant CPU maker for over a decade. ARM's history is another example.
- jnovek 7mo agoI don’t have a micro architecture background so I apologize if this is obvious — What do power and speed mean in this context?
- McP 7mo agoPower - how many Watts does it need? Speed - how quickly can it perform operations?
- wmf 7mo agoYou can get low power with a simple design at a low clock. This definitely will not help achieve high performance later.
- weebull 7mo agoClock rate isn't the only factor. A design can be power hungry at a low clock rate if designed badly, and if it it is... you're never getting that think running fast.
- unethical_ban 7mo agoOne could say "Optimize for efficiency first, then performance".
- jauntywundrkind 7mo ago
- rwmj 7mo agoMarcin is working with us on RISC-V enablement for Fedora and RHEL, he's well aware of the problem with current implementations. We're hopeful that this'll be pretty much resolved by the end of the year.
- cogman10 7mo ago> AND the software with no architecture-specific optimisations The optimizations that'd be applied to ARM and MIPS would be equally applicable to RISC-V. I do not believe this is a lack of software optimization issue. We are well past the days where hand written assembly gives much benefit, and modern compilers like gcc and llvm do nearly identical work right up until it comes to instruction emissions (including determining where SIMD instructions could be placed). Unless these chips have very very weird performance characteristics (like the weirdness around x86's lea instruction being used for arithmetic) there's just not going to be a lot of missed heuristics.
- hrmtst93837 7mo ago[flagged]
- cogman10 7mo agoWhile true, it's typically not going to be impactful on system performance. There's a reason, for example, why the linux distros all target a generic x86 architecture rather than a specific architecture.
- spockz 7mo agoNot all. CachyOS has specific builds for v3, v4, and AMD Zen4/5: https://wiki.cachyos.org/features/optimized_repos/ https://wiki.cachyos.org/features/optimized_repos/
- adrian_b 7mo agoSome applications may target a generic x86 architecture without any impact on performance. However, other applications which must do cryptographic operations, audio/video processing, scientific/technical/engineering computing, etc. may have wildly different performances when compiled for different x86-64 ISA versions, for which dedicated assembly-language functions exist.
- cogman10 7mo ago
- fidotron 7mo ago> RISC-V will get there, eventually. Not trolling: I legitimately don't see why this is assumed to be true. It is one of those things that is true only once it has been achieved. Otherwise we would be able to create super high performance Sparc or SuperH processors, and we don't. As you note, Arm once was fast, then slow, then fast. RISC-V has never actually been fast. It has enabled surprisingly good implementations by small numbers of people, but competing at the high end (mobile, desktop or server) it is not.
- gt0 7mo agoI don't think anybody suggests Oracle couldn't make faster SPARC processors, it's just that development of SPARC ended almost 10 years ago. At the time SPARC was abandoned, it was very competitive.
- twoodfin 7mo agoIn single-threaded performance? That’s not how I remember it: Sun was pushing parallel throughput over everything else, with designs like the T-Series & Rock.
- gt0 7mo agoPerhaps not single thread, but Rock was a dead end a while before Oracle pulled the plug, and Sun/Oracle's core market of course was always servers not workstations. We used Niagara machines at my work around the T2 era, a long time ago, but they were very competitive if you could saturate the cores and had the RAM to back it up.
- twoodfin 7mo agoSure, my work got a few of the Niagaras too and they were tremendous build machines for Solaris software. But if you’re judging an ISA by performance scalability, you generally want to look at single-threaded performance.
- icedchai 7mo ago
- newpavlov 7mo agoIn some cases RISC-V ISA spec is definitely the one to blame: 1) https://github.com/llvm/llvm-project/issues/150263 https://github.com/llvm/llvm-project/issues/150263 2) https://github.com/llvm/llvm-project/issues/141488 https://github.com/llvm/llvm-project/issues/141488 Another example is hard-coded 4 KiB page size which effectively kneecaps ISA when compared against ARM.
- adastra22 7mo agoAlso the bit manipulation extension wasn't part of the core. So things like bit rotation is slow for no good reason, if you want portable code. Why? Who knows.
- fidotron 7mo agoThe fact the Hazard3 designer ended up creating an extension to resolve related oddities was kind of astonishing. Why did it fall to them to do it? Impressive that he did, but it shouldn't have been necessary.
- rllj 7mo agoWhich extension is that?
- mjmas 7mo agoAn extension he calls Xh3bextm. For extracting multiple bits from bitfields. https://wren.wtf/hazard3/doc/#extension-xh3bextm-section https://wren.wtf/hazard3/doc/#extension-xh3bextm-section There are also four other custom extensions implemented.
- wren6991 7mo agoThis extension wasn't strictly necessary but it makes decode of Arm instructions faster in the bootrom's Arm emulator.
- 7mo ago
- Dwedit 7mo agoThere's the ARM video from LowSpecGamer, where they talk about how they forgot to connect power to the chip, and it was still executing code anyway. According to Steve Furber, the chip was accidentally being powered from the protection diodes alone. So ARM was incredibly power efficient from the very beginning.
- bsder 7mo ago> Don't blame the ISA - blame the silicon implementations That's true, but tautological. The issue is that the RISC-V core is the easy part of the problem, and nobody seems to even be able to generate a chip that gets that right without weirdness and quirks. The more fundamental technical problem is that things like the cache organization and DDR interface and PCI interface and ... cannot just be synthesized. They require analog/RF VLSI designers doing things like clock forwarding and signal integrity analysis. If you get them wrong, your performance tanks, and, so far, everybody has gotten them wrong in various ways. The business problem is the fact that everybody wants to be the "performance" RISC-V vendor, but nobody wants to be the "embedded" RISC-V vendor. This is a problem because practically anybody who is willing to cough up for a "performance" processor is almost completely insensitive to any cost premium that ARM demands. The embedded space is hugely sensitive to cost, but nobody is willing to step into it because that requires that you do icky ecosystem things like marketing, software, debugging tools, inventory distribution, etc. This leads to the US business problem which is the fact that everybody wants to be an IP vendor and nobody wants to ship a damn chip. Consequently, if I want actual RISC-V hardware, I'm stuck dealing with Chinese vendors of various levels of dodginess.
- crest 7mo agoRISC-V lacks a bunch of really useful relatively easy to implement instructions and most extensions are truly optional so you can't rely on them. That's the problem if you let a bunch of academics turn your ISA into a paper mill. In theory you can spend a lot of effort to make a flawed ISA perform, but it will be neither easy nor pretty e.g. real world Linux distros can't distribute optimised packages for every uarch from dual-issue in-order RV64GC to 8-wide OoO RV64 with all the bells and whistles. Only in (deeply) embedded systems can you retarget the toolchain and optimise for each damn architecture subset you encounter.
- userbinator 7mo agoARM was never a "speed demon"; it started out as a low power small-area core and clearly had more complexity and thought put into it than MIPS or RISC-V. Over a decade ago: https://news.ycombinator.com/item?id=8235120 https://news.ycombinator.com/item?id=8235120 RISC-V will get there, eventually. Strong doubt. Those of us who were around in the 90s might remember how much hype there was with MIPS.
- rbanffy 7mo agoI don’t think you remember, But the first Archimedes smoked the just-launched Compaq 386s with a dedicated 387 coprocessor. It was not designed to be one, but it ended up being surprisingly fast.
- izacus 7mo agoIf you make a spec that the wider industry cannot effectively implement into quality products, it's the spec that's wrong. And that's true for anything - whether it's RISC-V, ipv6, Matter, USB-C and so on. That's what makes writing specs hard - you need people who understand implementation challenges at the table, not dreaming architects and academics.
- leni536 7mo agoIs cross compilation out of the question?
- IshKebab 7mo agoIt's usually an enormous pain to set up. QEMU is probably the best option.
- sofixa 7mo agoDepends on the language, it's pretty trivial with Go.
- STKFLT 7mo agoMaybe there are issues I'm not aware of but using dockcross has made cross-compilation quite easy in my experience. https://github.com/dockcross/dockcross https://github.com/dockcross/dockcross
- mort96 7mo agoHow does it handle .so version differences and glibc version differences between the container and the target system?
- pantalaimon 7mo agoT2 manages to do it https://t2linux.com/ https://t2linux.com/
- 7mo ago
- lifis 7mo agoOr they could fix cross compilation and then compile it on a normal x86_64 server
- mort96 7mo agoFixing cross compilation is a huge undertaking. So much software needs to be patched to be properly cross-compilable.
- IshKebab 7mo agoYeah it's a few years behind ARM, but not that many. Imagine trying to compile this on ARM 10 years ago. It would be similarly painful.
- hackerInnen 7mo agoThis. While I doubt that there will be a good (whatever that means) desktop risc-v CPU anytime soon, I do think that it will eventually catch up in embedded systems and special applications. Maybe even high core count servers. It just takes time, people who believe in it and tons of money. Will see where the journey goes, but I am a big risc-v believer
- NetMageSCW 7mo agoWhy? They have yet to show anything to believe in except perhaps the embedded space.
- LeFantome 7mo agoYou think Meta bought Rivos to work on embedded? You think the Alibaba C930 CPU is for embedded? 15 SPECint2006 / GHz Or that the Tenstorrent Ascaclon will be? 18 SPECint2006 / GHz Even the SpacemiT K3 has better AI performance than an Apple Silicon M4. And RISC-V chips released this year are 2-4 times faster than last year. RISC-V is not the fastest ISA but it is improving the fastest. With so many companies backing RISC-V, why would I bet against it?
- kllrnohj 7mo ago> Imagine trying to compile this on ARM 10 years ago Cortex A57 is 14 years old and is significantly faster than the 9 year old Cortex A55 these RISC-V cores are being compared against. So yes it's many years behind. Many, many years.
- LeFantome 7mo agoSpacemiT K3 is on par with Rockchip RK3588. So, about 4 years behind ARM. Tenstorrent Atlantis (first Ascalon silicon) should ship in Q2/Q3 and be twice as fast. About as fast as Ryzen5. So, about 5 years behind AMD. But even the K3 has faster AI than Apple Silicon or Qualcomm X Elite. Current trend-lines suggest ARM64 and RISC-V performance parity before 2030.
- rbalint 7mo agoIf the builds are slow, build accelerators can help a lot. Ccache would work for sure and there is also firebuild, that can accelerate the linker phase and many other tools in builds.
- brcmthrowaway 7mo agoWhy is it slow? I thought we have Rivos chips
- Joel_Mckay 7mo agoAny new hardware lags in compiler optimizations. i. llvm presentation can thrash caches if setup wrong (given the plethora of RISC-V fragmented versions, most compilers won't cover every vanity silicon.) ii. gcc is also "slow" in general, but is predictable/reliable iii. emulation is always slower than kvm in qemu It may seem silly, but I'd try a gcc build with -O0 flag, and a toy unit test with -S to see if the ASM is actually foobar. One may have to force the -mtune=boom flag to narrow your search. Best regards =3
- Levitating 7mo agoThis is why felix has been building the risc-v archlinux repositories[1] using the Milk-V Pioneer. I think the ban of SOPHGO is part to blame for the slow development.[2] They had the most performant and interesting SOCs. I had a bunch of pre-orders for the Milk-V Oasis before it was cancelled. It was supposed to come out a while ago, using the SG2380, supposedly much more performant than the Milk-V Titan mentioned in the article (which still isn't out). It was also SOPHGO's SOCs that powered the crazy cheap/performant/versatile Milk-V DUO boards. They have the ability to switch ARM/RISC-V architecture. [1]: https://archriscv.felixc.at/ https://archriscv.felixc.at/ [2]: https://www.tomshardware.com/tech-industry/artificial-intelligence/us-govt-set-to-ban-huawei-intermediary-sophgo-over-ai-chip-supplies-partnership-skirted-us-chip-sanctions https://www.tomshardware.com/tech-industry/artificial-intell...
- 15155 7mo agoCan you articulate why you think this ban impacted anything and what you think the ban applies to?
- Levitating 7mo agoI won't pretend to understand the geo-politics or rulings. What I do know is since the ban, all ongoing products featuring SOPHGO SOCs were cancelled, and I haven't seen any products featuring them since. The SOPHGO forums have also closed down. The Milk-V Oasis would have had 16 cores (SG2380 w/ SiFive P670), it was replaced by the Milk-V Megrez with just 4 cores (SiFive P550) for around the same price. The new Milk-V Titan has only 8. We're slowly catching up, but the performance is now one or two years behind what it could've been. The SG2380 would've been the first desktop ready RISC-V SOC at an affordable price. I think it's still the only SOC made that used the SiFive P670 core.
- throwaway27448 7mo ago[flagged]
- ephou7 7mo agoUlrich Drepper, Lennart Poettering, this clown. Red Hat seems to have a skill of hiring savants with high technical and low social aptitude.
- primis 7mo agoHey! I get this is a throwaway account so you might not answer, but I really, really don't like opening an article and having the first thing I see in a thread be someone calling the author a slur. There are ways of expressing insult without bringing intellectual disabilities into the mix.
- dmit 7mo agoFor future readers: throwaway27448's comment used to say something completely different, featuring the r-slur, and then immediately edited.
- throwaway27448 7mo ago[flagged]
- notenlish 7mo agoCan you explain why you think the author is stupid.
- deleted 7mo ago[deleted]
- throwaway27448 7mo ago[flagged]
- andrepd 7mo agoThere's zero mention of hardware specs or cost beyond architecture and core counts... What is the purpose of this post? Anyway, it's hardly surprising that a young ISA with not a 1/1000th of the investment of x86 or ARM has slower chips than them x)
- kashyapc 7mo agoOn benchmarks, for more precision details, I recommend the RISC-V Vector (RVV) benchmarks[1], maintained by Olaf Bernsten. He only covers the Vector stuff, but with great depth. [1] https://camel-cdr.github.io/rvv-bench-results/ https://camel-cdr.github.io/rvv-bench-results/
- theodric 7mo ago[flagged]
- sltkr 7mo agoAre you sure you are comparing apples with apples here? The fact that i686 is 14% faster than x86_64 is a little suspicious, because usually the same software runs _faster_ on x86_64 (despite the increased memory use) thanks to a larger register set, an optimized ABI, and more vector instructions. Of course, if you are compiling an i686 binary on i686, and an x86_64 binary on x86_64, then the compilers aren't really doing the same work, since their output is different. I'm not a compiler expert, but I could imagine that compiling x86_64 binaries is intrinsically slower than for i686 for a variety of reasons. For example, x86_64 is mostly a superset of i686, so a compiler has way more instructions to consider, including potential optimizations using e.g. SIMD instructions that don't exist on i686 at all. Or a compiler might assume a larger instruction cache size, by default, and do more unrolling or inlining when compiling for x86_64. And so on. In that case, compiling on x86_64 is slower not because the hardware is bad but because the compiler does more work. Perhaps something similar is happening on RISC-V.
- srott 7mo agoCouldn’t be caused by a slower compiler? Fe. What would be a difference when cross compiling same code to aarch64 vs risc-v?
- yogthos 7mo agothere are projects for making high performance RISC-V chips like this one https://github.com/OpenXiangShan/XiangShan https://github.com/OpenXiangShan/XiangShan
- classichasclass 7mo agoOK, I'll bite. If this is a truly competitive core - I don't claim enough personal expertise to judge - does anyone fab and sell it? There should be a business case if it is.
- luyu_wu 7mo agoIf I remember correctly,it was taped out by some company as some embedded core in a GPU? I guess that may be the true use case for 'Open-Source' cores. That being said, the advertised SPEC2007 scores are close to a M1 in IPC.
- kashyapc 7mo agoArm had 40 years to be where it is today. RISC-V is 15 years old. Some more patience is warranted. Assuming they will keep their word, later this year Tenstorrent is supposed to ship their RVA23-based server development platform[1]. They announced[2] it at the last year's NA RISC-V Summit. Let's see. The ball is in the court of hardware vendors to cook some high-end silicon. [1] https://tenstorrent.com/ip/risc-v-cpu https://tenstorrent.com/ip/risc-v-cpu [2] https://static.sched.com/hosted_files/riscvsummit2025/e2/Unleash%20your%20RISC-V%20Future%20with%20Tenstorrent%E2%80%99s%20High%20Performance%20Ascalon%20RISC-V%20Processor%20-%20Now%20Available%21%20-%20Troy%20Jones%2C%20Tenstorrent.pdf https://static.sched.com/hosted_files/riscvsummit2025/e2/Unl...
- userbinator 7mo agoMIPS, which RISC-V is closely modeled after, is also roughly 4 decades old and was massively hyped in the early 90s as well.
- kashyapc 7mo agoGreat point; I only know about MIPS legacy vaguely. As you imply, don't listen to the "hype-sters" but pay attention to what silicon is being produced.
- saati 7mo agoAarch64 is just 15 years old, and shares pretty much nothing with 32 bit arms apart from the name.
- mrbluecoat 7mo ago> Random mumblings of ARM developer ... RISC-V is sloooow Old news. See also: > Random mumblings of x86_64 developer ... ARM is sloooow
- throwa356262 7mo agoWhat kind or ancient arm hardware are they using here? On a related note, SoC companies needs to get their act together and start using the latest arm cores. Even the mid range cores of 1-2 years ago show a huge leap in performance: https://sbc.compare/56-raspberry-pi-500-plus-16gb/101-radxa-orion-o6n-32gb https://sbc.compare/56-raspberry-pi-500-plus-16gb/101-radxa-...
- orangeboats 7mo ago>What kind or ancient arm hardware are they using here? I think that's the point being made here. ARM in the 2000s was not known to be fast, now it is. RISC-V being slow isn't an inherent characteristic of the ISA, it only tells you about the quality of its implementations. And said implementations will only improve if corporations are throwing capitals at it (see: Apple, Qualcomm, etc.)
- throwa356262 7mo agoI think standard Arm cores are already plenty fast, the issue is the SoC vendors are still using cortex-A57 from 2015 instead of the new designs.
- orangeboats 7mo agoI am not talking about modern ARM though.
- saghm 7mo agoIf I'm reading their chart right, they have barely half as much memory for their RISC-V machine compared to any of the others? I don't know enough to know whether it's actually bottlenecked by memory, but it's a bit odd to claim it's slower, give those numbers, and not say anything about it. I'd hope they ruled that out as the source of the discrepancy, but it's hard to tell without confirmation.
- Levitating 7mo agoI think it's mentioned clearly in the article. > RISC-V builders have four or eight cores with 8, 16 or 32 GB of RAM (depending on a board) > The UltraRISC UR-DP1000 SoC, present on the Milk-V Titan motherboard should improve situation a bit (and can have 64 GB ram). RISC-V SOCs just typically don't support much ram. With the exception of the SG2042 which can take 128GB, but it's expensive, buggy and now old. So I am sure it's a combination of low ram and low clockspeeds.
- saghm 7mo agoThat sounds a lot less "RISC-V is slow" and more like "the most money I'm willing to spend on a RISC-V machine is low, but the more powerful ones may or not be as slow". I guess that doesn't make a particularly compelling headline.
- AceJohnny2 7mo agoThere was a Mastodon post some time back (~1y?) where someone realized that the fastest RISC-V hardware they could get was still slower than running it on QEMU. That's not how it usually works :\ RISC-V is certainly spreading across niches, but performant computing is not one of them. Edit: lol the author mentions the same! Perhaps they were the source of the original Mastodon post I'm thinking of.
- Levitating 7mo agoThe Milk-V Pioneer breaks that barrier, it's expensive though. And the risc-v architecture used is now old, the company that developed is was sanctioned by the US and is now dead.
- mkj 7mo agoDoes that page even say which RISC-V CPUs are being used that are slow? I couldn't see it, which seems a bit of pointless complaining.
- Levitating 7mo ago> RISC-V builders have four or eight cores with 8, 16 or 32 GB of RAM (depending on a board). Which boards are used specifically should not matter much. There's not much available. Except for the Milk-V Pioneer, which has 64 cores and 128GB ram. But that's an older architecture and it's expensive.
- echoangle 7mo agoIs there a simple explanation why RISC-V software has to be built on a RISC-V system? Why is it so hard for compilers to compile for a different architecture? The general structure of the target architecture lives inside the compiler code and isn’t generated by introspecting the current system, right?
- boredatoms 7mo agoUnder specified build dependencies that use libraries/config on your host OS rather than the target system You can solve this on a per language basis, but the C/C++ ecosystem is messy. So people use VMs or real hardware of the target arch to not have to think about it
- anarazel 7mo agoCross building of possible, but it's rather useful to be able to test the software you just built... And often enough, tests take more resources than the build.
- flowerthoughts 7mo agoOld compilers tended to make it a compile-time switch which backends were included, probably because backends were "huge", so they were left out. (The insn lookup table in GCC took ages to generate and compile.) And of course all development environments running on Windows assumed x86 was the only architecture. With LLVM existing, cross-compiling is not a problem anymore, but it means you can't run tests without an emulator. So it might just be easier to do it all on the target machine.
- AnssiH 7mo agoThe cross-compiler part itself is easy, but getting all the build scripting of tens of thousands of Fedora packages to work perfectly for cross-compiling would be a lot of work. There are lots of small issues (libraries or headers not being found, wrong libraries or headers being found, build scripts trying to run the binaries they just built, wrong compiler being used, wrong flags being used, etc.) when trying to cross-compile arbitrary software. All fixable (cross-compiling entire distributions is a thing), but a lot of work and an extra maintenance burden.
- devl547 7mo agoIs it RISC-V or bloated software full of layered abstractions?
- deleted 7mo ago[deleted]
- utopiah 7mo agoFWIW checkout dockcross/linux-riscv32 and dockcross/linux-riscv64 if compilation itself is your problem. I setup a CopyParty server on a headless RISC-V SBC and was a breeze. Just get the packets, do the thing, move on. Obviously depends on your need but maybe you're not using the right workflow and blame the tools instead.
- aa-jv 7mo agoI don't care as long as it keeps my soldering iron hot.
- sylware 7mo agoThe current hardware used is self-hosting mini-server grade, and certainly not on the latest silicon process. "Slow" is expected. It is not the ISA, but the implementations and those horrible SDKs which needs to be adjusted for RISC-V (actually any new ISA). RISC-V needs extremely performant implementations, that on the best silicon process, until then RISC-V _will be_ "slow". Not to mention, RISC-V is 'standard ISA': assembly writted software is more than appropriate in many cases.
- kashyapc 7mo agoA couple of corrections (the blog-post is by a colleague, but I'm not speaking for Marcin! :)) First, we do have a recent 'binutils' build[1] with test-suites in 67 minutes (it was on Milk-V "Megrez") in the Fedora RISC-V build system. This is a non-trivial improvement over the 143-minute build time reported in the blog. Second, the current fastest development machine is not Banana Pi BPI-F3. If we consider what is reasonably accessible today, it is SiFive "HiFive P550" (P550 for short) and an upcoming UltraRISC "DP1000", we have access to an eval board. And as noted elsewhere in this thread, in "several months" some RVA23-based machines should be available. (RVA23 == the latest ISA spec). FWIW, our FOSDEM talk from earlier this year, "Fedora on RISC-V: state of the arch"[1], gives an overview of the hardware situation. It also has a couple of related poorman's benchmarks (an 'xz' compression test and a 'binutils' build without the test-suite on the above two boards -- that's what I could manage with the time I had). Edit: Marcin's RISC-V test was done on StarFive "Vision Five 2". This small board has its strengths (upstreamed drivers), but it is not known for its speed! [1] https://riscv-koji.fedoraproject.org/koji/taskinfo?taskID=91687 https://riscv-koji.fedoraproject.org/koji/taskinfo?taskID=91... [2] Slides: https://fosdem.org/2026/events/attachments/SQGLW7-fedora-on-riscv/slides/266759/fedora-ri_x3tr93d.pdf https://fosdem.org/2026/events/attachments/SQGLW7-fedora-on-...
- brucehoult 7mo ago> VisionFive 2 It's a good solid reliable board, but over three years old at this point (in a fast-moving industry) and the maximum 8 GB RAM is quite challenging for some builds. Binutils is fine, but on recent versions of gcc it wants to link four binaries at the same time, with each link using 4 GB RAM. I've found this fails on my 16 GB P550 Megrez with swap disabled, but works quickly and uses maybe 50 or 100 MB of swap if I enable it. On the VisionFive 2 you'd need to use `-j1` (or `-j2` with swap enabled) which will nearly double or quadruple the build time. Or use a better linker than `ld`. At least the LLVM build system lets you set the number of parallel link jobs separately to the number of C/C++ jobs.
- kashyapc 7mo ago> I've found this fails on my 16 GB P550 Megrez with swap disabled but works quickly and uses maybe 50 or 100 MB of swap if I enable it. I see, I don't have a Megrez at my desk, only in the build system. I only have P550 as my "workhorse". PS: I made a typo above - the P550 I was referring to was the SiFive "HiFive Premier P550". But based on your HN profile text, you must've guessed it as much :)
- cesaref 7mo agoJust out of interest, why aren't they cross compiling RISC-V? I thought that was common practice when targeting lower performing hardware. It seems odd to me that the build cycle on the target hardware is a metric that matters.
- kashyapc 7mo agoPlease skim the thread :) We've already discussed it twice. Fedora "mandates" native builds. Build time on target hardware matters when you're re-building an entire Linux distribution (25000+ packages) every six months.
- cesaref 7mo agoI failed to find this on my skim, my bad :( Interesting that it's mandated as native - i'm really not sure the logic behind this (i've worked in the embedded world where such stuff is not only normal, but the only choice). I'll do some digging and see if I can find the thought process behind this.
- poulpy123 7mo agoIs it slow because of the inherent design or because it's recent and not as optimised as x86 or arm ?
- haerwu 7mo agoI updated blog post after reading comments from Matrix/Slack/Phoronix/HN/Lobster/etc. places. - mentioned which board had 143 minutes, added info about time on Milk-V Megrez board - added section 'what we need hw-wise for being in fedora' - added link to my desktop post to point that it is aarch64, not x86-64 - wording around qemu to show that I use it locally only
- titzer 7mo ago> ... I can build the “llvm15” package in about 4 hours. Compare that to 10.5 hours on a Banana Pi BPI-F3 builder (it may be quicker on a P550 one). That's....slow. What a huge pile of bloat.
- Steinmark 7mo ago[dead]
- shmerl 7mo agoWhy not cross compile in such case on better hardware? Then run tests on the native one.
- 0verkilled 7mo agoUnrelated to the post's point but: Why does x86 build faster than x86_64? Presumably they used the same exact hardware, or at least the exact same number of cores and memory, yet the build time is more than 10% faster in x86. Is there some sort of overhead for x86_64 that I'm not seeing?
- rivetfasten 7mo agoThanks for the post! Question: While you would want any official arch built natively, maybe an interim stage of emulated vm builds for wip/development/unsupported architectures would still be preferable in this case? Comparing the tradeoffs: * Packages disabled and not built because of long build times. * Packages built and automated tests run on inaccurately emulated vms (NOT cross compiled). Users can test. It might be broken. It's an experimental arch, maybe the build cluster could be experimental too?
- lifeline82 7mo ago[dead]
- rurban 7mo agoWindows is still much slower.
- LeFantome 7mo agoThis is article is being discussed on another forum where kernel build times are being compared for different RISC-V hardware. The conclusion there was that, if a BananaPi-F3 is taking 143 minutes to compile binutils, the SpacemiT K3 will buld it in 36 minutes using its X100 cores (half its cores). That is the same as the time he quotes for the unidentified Aarch64 hardware. Which makes this a pretty funny article. I do not have a K3 to confrim. I am hoping to pick one up when it becomes more widely available next month.
- LeFantome 7mo agoI am going to make a wild guess here. The reason that he does not tell us what hardware he is using is because none of these times are for a single system building binutils. I think he is using a mix of systems and then doing some kind of averaging to tell us what a individual system would look like. For some kind of hardware, all the systems they have would be the fastest that architecture offers, like with i686 I expect. While others are going to be a mix of old and new, like x86-64. For RISC-V, the latest gen hardware is about as fast as the numbers he quotes for Aarch64. To be clear, the fastest ARM is still faster than the fastest RISC-V. But the numbers he quotes make no sense for something like a SpacemiT K3. But if you are using RISC-V systems from two years ago in your build cluster, they will as he says be "Sloooow". But that shows how fast RISC-V is improving. It makes no sense to publish this article now. At least, he should reveal what hardware he is talking about. His chart makes no sense (for most of the platforms).