28 ms·
“Risc V greatly underperforms”
- deleted 5y ago[deleted]
- jpfr 5y agoThe idea is to use the compressed instruction extension. Then two adjacent instructions can be handled like a single “fat” instruction with a special case implementation. That allows more flexibility for CPU designs to optimize transistor count vs speed vs energy consumption. This guy clearly did not look at the stated rationale for the design decisions of RISC-V.
- theresistor 5y agoCompressed instructions and macro-fusion aren't magical solutions. It's not always possible to convince the compiler to generate the magical sequence required, and it actually makes high-performance implementations (wide superscalar) more difficult thanks to the variable width decoding. Beyond that, compressed instructions are not a 1:1 substitute for more complex instructions, because a pair of compressed instructions cannot have any fields that cross the 16-bit boundary. This means you can't recover things like larger load/store offsets. Additionally, you can't discard architectural state changes due to the first instruction. If you want to fuse an address computation with a load, you still have to write the new address to the register destination of the address computation. If you want to perform clever fusion for carry propagation, you still have to perform all of the GPR writes. This is work that a more complex instruction simply wouldn't have to perform, and again it complicates a high performance implementation.
- jpfr 5y agoIn the context of gmp, people write architecture-specific assembly for the inner loop anyway. Besides that, you raise good points on sources of complexity. I’m waiting for the benchmarks once such developments have been incorporated. Everything else is guesswork.
- throwaway81523 5y agoIf they didn't implement those benchmarks (at least in simulation, like they benchmarked everything else) before releasing the spec, then they have nothing but handwaving and wishful thinking in saying this issue can be solved by op fusion. The reality is that they optimized for 1980s-style C programming without noticing that this isn't the 1980s any more.
- panick21_ 5y agoPart of the idea is to create standard ways to do certain things and then hope compiler writers generation code according to that. That will allow more chip designers to take advantage of those if they want to. They spent a lot of time and effort on making sure the decoding pretty good and useful for high performance implementations. RISC-V is designed for very small and very large system. At some point some tradeoffs need to be made but these are very reasonable and most of the time no a huge problem. For the really specialized cases where you simply can't live with those extra instruction, those will be added to the standard and then some profiles will include them and others not. If those instructions are really as vital as those that want them claim, they will find their way into many profiles. Saying RISC-V is 'terrible' because of those choices is not fair way of evaluating it.
- userbinator 5y agoRISC-V is designed for very small and very large system That's exactly the problem --- there is no one-size-fits-all when it comes to instruction set design.
- panick21_ 5y agoThere is a trade-off but there is overall far more value in having it be unified. The trade-offs are mostly very small or non existent once you consider the standard extensions that different use cases will have. Overall having a unified open instruction set is far better then hand designing many different instruction sets just to get marginal improvement. Some really extreme application might require that, but for the most part the whole indsutry could do just fine with RISC-V. Both on the low and on the high end, and in fact better then most of the alternative all things considered. If integer checking is really the be all end all and without it RISC-V can not be successful without it, it will be added and it will be pulled into all the profiles. If it is not actually that relevant then it wont. If it is very useful for some verticals and not others, it will be in those profiles and not in others.
- ddingus 5y ago>overall far more value in having it be unified. >[...] >If integer checking is really the be all end all and without it RISC-V can not be successful without it, it will be added and it will be pulled into all the profiles. If it is not actually that relevant then it wont. If it is very useful for some verticals and not others, it will be in those profiles and not in others. So which is it? Unified or something else?
- audunw 5y ago> and it actually makes high-performance implementations (wide superscalar) more difficult thanks to the variable width decoding. More difficult than x86? We're talking about a damn simple variable width decoding here. I could imagine RISC-V with C extension being more tricky than 64-bit ARM. Maybe. > and again it complicates a high performance implementation. But so much of the rationale behind the design of RISC-V is to simplify high performance implementation in other ways. So the big question is what the net effect is. The other big question is if extensions will be added to optimise for desktop/server workloads by the time RISC-V CPUs penetrate that market significantly.
- imtringued 5y agoLet's assume you are right. In 5 years the organization behind RISC-V apologizes and introduces a "bignum" extension. That doesn't sound too bad.
- socialdemocrat 5y agoI don’t see why offsets larger than 16-bit are important. Are you implying that most fusion candidate pairs would need this? In tight inner loops why would you need large offsets? Of course you discard architectural state changes in fusion. If I have a bunch of instructions which end up reading from memory into register x10, then I can fuse with all previous instructions which wrote into x10, as their results get clobbered anyway. Disclaimer: I may have misunderstood the point you made. However you don’t seem to make it clear how fusion is bad for performance. What performance tricks are you giving up by doing fusion?
- okl 5y agoSweet spot seems to be 16-bit instructions with 32/64-bit registers. With 64-bit registers you need some clever way to load your immediates, e.g., like the shift/offset in ARM instructions.
- msbarnett 5y agoHe literally addressed this, albeit obliquely, in the message > I have heard that Risc V proponents say that these problems are known and could be fixed by having the hardware fuse dependent instructions. Perhaps that could lessen the instruction set shortcomings, but will it fix the 3x worse performance for cases like the one outlined here? Macro-fusion can to some extent offset the weak instruction set, but you're never going to get a multiple integer multiplier speedup out of it given the complexity of inter-op architectural state changes that have to be preserved, and instruction boundary limitations involved; it's never going to offset a 3x blowup in instruction count in a tight loop.
- socialdemocrat 5y agoFusing 3 instructions is not unusual, those could also have been compressed. Thus you have no more microcode to execute and only 50% more cache usage rather than 300%
- alerighi 5y agoEven if you do so, the program size is still bigger, and it consumes more disk, RAM and most importantly cache space. Wasting cache for having multiple instructions when on another architecture it's done by only one doesn't make particular sense to me. Also, it's said that x86 is bad because the instructions are then reorganized and translated inside the CPU. But it seems that you are proposing the same, the CPU that preprocessed the instructions and fuses some into a single one (the opposite that x86 does). Ad that point, it seems to me that what x86 does makes more sense: have a ton of instruction (and thus smaller programs and thus more code that can fit in cache) and split them, rather than having a ton of instructions (and waste cache space) for then the CPU to combine them into a single one (a thing that a compiler can also do).
- Buttons840 5y agoHow many cache misses are for program instructions, versus data misses?
- anarazel 5y agoIME icache misses are a frequent bottleneck. There's plenty code where all the time is spent in one tight inner loop and thus the icache is not a constraint, but there's also a lot of cases with a much flatter profile. Where icache misses suddenly become a serious constraint.
- imtringued 5y agoObtaining the carry bit does not involve branches though. Overflow checking probably does.
- alerighi 5y agoDepends on the application. But even if they are few, it's not a good reason to have them just for having a nice instruction set, that if you are not writing assembly by hand (and nobody does these days) doesn't give you any benefit. Also don't reason with the desktop or server use case in mind, where you have TB of disk and code size doesn't matter. RISC-V is meant to be used also for embedded systems (in fact their use nowadays is only for these systems), where usually code size matter more than performance (i.e. you typically compile with -Os). In these situations more instructions means more flash space wasted, meaning you can fit less code.
- Taniwha 5y agoSo this is one tiny corner of the ISA, not something that makes ALL instruction sequences longer - essentially RISCV has no condition codes (they're a bit of an architectural nightmare for everyone doing any more than the simplest CPUs, they make every instruction potentially have dependencies or anti-dependencies with every other). It's a trade off - and the one that's been made is one that makes it possible to make ALL instructions a little faster at the expense of one particular case that isn't used much - that's how you do computer architecture, you look at the whole, not just one particular case RISCV also specifies a 128-bit variant that is of course FASTER than these examples
- wbl 5y agoYou can opt in to generating and propagating conditions and rename the predicates as well.
- theresistor 5y agoThis isn't an isolated case. RISC-V makes the same basic tradeoff (simplicity above all else) across the board. You can see this in the (lack of) addressing modes, compare-and-branch, etc. Where this really bites you is in workloads dominated by tight loops (image processing, cryptography, HPC, etc). While a microarchitecture may be more efficient thanks to simpler instructions (ignoring the added complexity of compressed instructions and macro-fusion, the usual suggested fixes...), it's not going to be 2-3x faster, so it's never going to compensate for a 2-3x larger inner loop.
- rstuart4133 5y ago> it's not going to be 2-3x faster, so it's never going to compensate for a 2-3x larger inner loop. As someone else who replied said, I'm not a CPU architect, just software that works close to the metal. That means I pay attention to compiler output. What you say is true in the very early says: compilers did indeed use the x86's addressing modes in all sorts of odd ways to squeeze as many calculations as possible into as few bytes as possible. Then it went in the reverse direction. You started seeing compilers emitting long series of simple instructions instead, seemingly deliberately avoiding those complex addressing modes. And now it's swung back again - I'm the complier using addressing modes to shift plus a couple of adds in one instruction is common again. I presume all these shifts were driven by speed of the resulting code. I have no idea why one method was faster than the other - but clearly there is no hard and fast rule operating here. For some internal x86's implementations using complex addressing modes was a win. On some, for exactly the same instruction set, it wasn't. There is no cut and dried "best" way of doing it, rather it varies as the transistor and power budget changes. One thing we do know about RISC-V is it is intended to cover a _lot_ transistor and power budgets. Where it's used now (low power / low transistor) is turned out their design decisions have turned out _very_ well, far better than x86. More fascinatingly to me, today the biggest speed ups compilers get for super scalar arch's has nothing to do with the addressing modes so much attention is being focused on here. It comes from avoiding conditional jumps. The compilers will often emit code that evaluates both paths of the computation (thus burning 50% more ALU time on a calculating a result that will never be used), then choose the result they want with a cmov. In extreme cases, I've seen doing that sort of thing gain them a factor of 10, which is far more than playing tiddly winks with addressing modes will get you. I have no idea how that will pan out for RISC-V. I don't think any one has done a super scalar implementation of it yet(?) But in the non-super scalar implementations the RISC-V instruction set choices have worked out very well so far. And when someone does do a super scalar implementation (and I'm sure there will be a lot of different implementations over time), it seems very possible x86's learnings on addressing mode use will be yesterdays news.
- aappleby 5y agoThe author seems to be assuming that the designers have never thought about this corner case.
- sanxiyn 5y agoNo, the author is arguing this is not a corner case but a central? case. I tend to agree.
- sosodev 5y agoWhy do these half baked slam pieces always make it to the top of HN?
- chillingeffect 5y agoif for no other reason than to quickly formulate counterarguments. Next time at some meeting or other get together, if someone pipes up with an anti-RISC comment, most people won't be able to quickly refute it. But having had this discussion here, we're inocculated and able to respond with intelligence and experience.
- okl 5y agoThat sounds like you make up your mind first, then look for arguments that support your position. I'd rather see the arguments before I come to conclusions.
- jgilias 5y agoMany people upvote things not necessarily because they agree with them, but rather to bump it in hopes that someone with good insights will chime in in the comments section. This especially applies to potentially controversial things.
- bob1029 5y agoI think the reason is that it ultimately encourages deep and thoughtful conversation. If nothing controversial was ever proposed, the motivation for participating and "proving others wrong" is lessened. It might not be the healthiest way, but I certainly find myself putting a lot more thought into my comments if its a contrary point or in some broader controversial context. Overall, I feel HN is most fun when a lot of people are in disagreement but also operating in good faith.
- boibombeiro 5y agoStanding ground, specially when we are wrong, helps to learn a lot more about the subject.
- nynx 5y agoIf this really is an issue, I imagine risc-v could easily get an extension for adding/subtracting/etc simd vectors together in a way that would expand to the capabilities of underlying processor without requiring hardcoding.
- samstave 5y agoSo a few decades ago.. I knew a guy who was one of the chief designers of RISC procs at MIPs Yeah - he was addicted to prostitutes... This guy was doing amazing engineering work and we talked at length about designing a system for basically what became rack-mount trays But he was so distracted by his addiction to prostitutes...
- Symmetry 5y agoI think talking about ISAs as better or worse than one another is often a bad idea for the same reason that arguing about whether C or Python is better is a bad idea. Different ISAs are used for different purposes. We can point to some specific things as almost always being bad in the modern world like branch delay slots or the way the C preprocessor works but even then for widely employed languages or ISAs there was a point to it when it was created. RISC-V has a number of places it's employed where it makes an excellent fit. First of all academia. For an undergrad making building the netlist for their first processor or a grad student doing their first out of order processor RISC-V's simplicity is great for the pedagogical purpose. For a researcher trying to experiment with better branch prediction techniques having a standard high-ish performance open source design they can take and modify with their ideas is immensely helpful. And for many companies in the real world with their eyes on the bottom line like having an ISA where you can add instructions that happen to accelerate your own particular workload, where you can use a standard compiler framework outside your special assembly inner loops, and where you don't have to spend transistors on features you don't need. I'm not optimistic about RISC-V's widescale adoption as an application processor. If I were going to start designing an open source processor in that space I'd probably start with IBM's now open Power ISA. But there are so many more niches in the world than just that and RISC-V is already a success in some of them.
- okl 5y agoBranch delay slots are an artifact of a simple pipeline without speculation. There's nothing inherently "bad" about them.
- pm215 5y agoIf you're designing a single CPU that definitely has a simple pipeline, branch delay slots are maybe justifiable. If you're designing an architecture which you hope will eventually be used by many CPU designs which might have a variety of design approaches, then delay slots are pretty bad because every future CPU that isn't a simple non-speculating pipeline will have to do extra work to fake up the behaviour. This is an example of a general principle, which is that it's usually a mistake to let microarchitectural details leak into the architecture -- they quickly go stale and then both hw and sw have to carry the burden of them.
- okl 5y agoFew years ago, I designed my own ISA. In that time I investigated design decisions in lots of ISAs and compared them. There was nothing in the RISC-V instruction set that stood out to me, like for example, the SuperH instruction set, which is remarkably well designed. Edit: Don't get me wrong, I don't think RISC-V is "garbage" or anything like that. I just think it could have been better. But of course, most of an architecture's value comes from its ecosystem and the time spent optimizing and tailoring everything...
- AlotOfReading 5y agoMy memories of SuperH are a bit different. Yeah, it's cleaner than ARM, but the delay slots, hardware division, and the tiny register file among others made life unnecessarily difficult. A lot of those design decisions didn't hold up well over time.
- okl 5y agoInteresting! From which perspective? Implementing the ISA, compiler or applications? Did you write machine language or compiled?
- AlotOfReading 5y agoMainly system level and higher, but a bit of all three, I suppose. I was helping reverse engineer a customized SH chip and ended up implementing a small VM and optimized system libraries/utilities afterwards. Most of the time was spent in assembly, with some machine code and C on either side.
- okl 5y agoThanks for your insight.
- mlyle 5y agoIt's not really too unusual of a vantage point. SH was clean in some ways, but delay slots are annoying for anything that speculatively executes (which sets a bit of a performance ceiling without a lot of complexity), and more registers are generally better.
- Teknoman117 5y agoA bit of a computer history question: I have never looked at the ISA of the Alpha (referenced in post), but RISC V has always struck me as being nearly identical to (early) MIPS, just without the HI and LO registers for multiply results and the addition of variable length instruction support, even if the core ISA doesn't use them. MIPS didn't have a flag register either and depended on a dedicated zero register and slt instructions (set if less than)
- okl 5y agoI bet this article on RISC-V's genealogy is interesting for you: https://live-risc-v.pantheonsite.io/wp-content/uploads/2016/02/EECS-2016-6.pdf https://live-risc-v.pantheonsite.io/wp-content/uploads/2016/...
- cpeterso 5y agoAndrew Waterman's thesis ("Design of the RISC-V Instruction Set Architecture") has a very approachable comparison of RISC-V to MIPS, SPARC, Alpha, ARMv7, ARMv8, OpenRISC, and x86: https://www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-1.pdf https://www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-...
- dfox 5y agoThe no flags at all part is clearly inspired by Alpha including the rationale of flags being detrimental to OoO implementation. MIPS is classical RISC design that was not designed to be OoO-friendly at all and is simply designed for ease of straightforward pipelined implementation. The reason why it does not have flags probably simply comes down to the observation that you don't need flags for C.
- deleted 5y ago[deleted]
- userbinator 5y agoYes, that's exactly my thought every time it comes out; RISC-V is likely to displace MIPS everywhere performance doesn't matter, but it'll have a hard time competing with ARM or x86 on that.
- deleted 5y ago[deleted]
- CalChris 5y agoTL;DR My code snippet results in bloated code for RISC-V RV64I. I'm not sure how bloated it is. All of those instructions will compress [1]. [1] https://riscv.org/wp-content/uploads/2015/05/riscv-compressed-spec-v1.7.pdf https://riscv.org/wp-content/uploads/2015/05/riscv-compresse... It's slower on RISC-V but not a lot on a superscalar. The x86 and ARMv8 snippets have 2 cycles of latency. The RISC-V has 4 cycles of latency. 1. add t0, a4, a6 add t1, a5, a7 2. sltu t6, t0, a4 sltu t2, t1, a5 3. add t4, t1, t6 sltu t3, t4, t1 4. add t6, t2, t3 I'm not getting terrible from this.
- Koffiepoeder 5y agoCPU performance increases nowadays often are measured in single digit percentages because the margins became so thin. Doubling the cycles is a 100% increase. You can call that not so bloated, but I think many people would beg to differ. On the other hand I take this article with a grain of salt anyhow, since it only discusses a single example. I think we would need a lot more optimized assembly snippet comparisons to make meaningful conclusions (and even then there could be authored selection bias).
- snvzz 5y agoThe article's approach to arguing against RISC-V is fairly childish. >"here's this snippet, it takes more instructions on RISC-V, thus RISC-V bad" Is pretty much what it's saying. An actual argument about ISA design would weight the cost this has with the advantages of not having flags, provide a body of evidence and draw conclusions from it. But, of course, that would be much harder to do. What's comparatively easy and they should have done, however, is to read the ISA specification. Alongside the decisions that were made, there's a rationale to support it. Most of these choices, particularly so the ones often quoted in FUD as controversial or bad, have a wealth of papers, backed by plentiful evidence, behind them.
- snvzz 5y agoI don't think they even tried to read the ISA spec documents. If they did, they would have found that the rationale for most of these decisions is solid: Evidence was considered, all the factors were weighted, and decisions were made accordingly. But ultimately, the gist of their argument is this: >Any task will require more Risc V instructions that any contemporary instruction set. Which is easy to verify as utter nonsense. There's not even a need to look at the research, which shows RISC-V as the clear winner in code density. It is enough to grab any Linux distribution that supports RISC-V and look at the size of the binaries across architectures.
- robert_foss 5y agoSo how would you suggest re-writing their example in less than 6 instructions for RISC-V? X86/arm both have instructions that include the carry operation for long additions, and only require 2 instructions.
- vitno 5y agoAny != All. There is a difference between synthetic benchmarks and real world test cases.
- api 5y agoSo this person found a pathological case for the RISC-V instruction set?
- adrian_b 5y agoThis is not a pathological case, it is normal operation. A computer is supposed to compute, but the RISC-V ISA does not provide everything that is needed for all the kinds of computations that exist. The 2 most annoying missing features are the lack of support for multi-word operations, which are needed to compute with numbers larger than 64 bits, but also the lack of support for detecting overflow in the operations with standard-size integers. If you either want larger integers or safe computations with normal integers, the number of RISC-V instructions needed for implementation is very large compared to any other ISA. While there are people who do a lot of computations with large numbers, even the other users need such operations every day. Large number computations are needed at the establishment of any Internet connection, for the key exchange. For software developers, many compilers, e.g. gcc (which uses precisely libgmp), do computations with large numbers during compilation, for various kind of optimizations related to the handling of constants in the code, e.g. for sub-expression extraction or for operation complexity lowering. So every time when some project is compiled, libgmp or other equivalent library for large numbers might be used, like also every time when you click on a new link in a browser. So this case is not at all pathological, except in the vision of the RISC-V designers who omitted support for this case. That was a good decision for an ISA intended only for teaching or for embedded computers, but it becomes a bad decision when someone wants to use RISC-V outside those domains, e.g. for general-purpose personal computers.
- throwaway19937 5y agoTL;DR RISC-V doesn't have add with carry. I'm not a fan of the RISC-V design but the presence or absence of this instruction doesn't make it a terrible architecture.
- stephencanon 5y ago_For the purposes of implementing multi-word arithmetic_, which is Torbjörn's whole deal, it kind of does. (Also the actual post subject is "greatly underperforms").
- FullyFunctional 5y agoIt's meaningless to look at the code in absence an implementation and conclude anything about the performance. He doesn't know what the performance is. Having six instruction vs. two does not mean one is 3X faster than the other. It means nothing at all.
- stephencanon 5y agoWe know enough about the implementation of current RISC-V cores to conclude that they won't be remotely competitive on this one narrow (yet fairly high-impact for some workloads) task. Is it _possible_ to design a core that is competitive on this workload even when handicapped by a limited ISA? Yes, definitely. Have any RISC-V designers shown any interest in doing so yet? No.
- kayamon 5y ago"Gee no carry flag how will we cope?"
- yjftsjthsd-h 5y agoThe original title was "Risc V greatly underperforms", which seems like a far more defensible and less inflammatory claim than "Risc V is a terrible architecture", which was picked from the actual message but still isn't the title.
- gary_0 5y agoI almost skipped this thread because of the flamebait title. This is a debate over CPU instruction set performance details, nobody is going to die.
- yjftsjthsd-h 5y agoIn fairness, this is Hacker News; flame wars^w^w respectful but intense debate over editors, operating systems, and, yes, ISA details, is somewhat expected. (Although, yes, I'm not sure that I would get too worked up about this particular detail; even if the stated claim is 100% true and unmitigated, it means some kinds of code will have potentially bigger binaries. I understand a math library person caring, I don't think I care.)
- rbanffy 5y ago> I understand a math library person caring, I don't think I care. Not wasting much sleep on this one. Not sure there's anything on the spec that stops implementations from recognizing the two instructions and fuse them into a single atomic operation for the backends to deal with. It'll occupy more space in the L1 cache, but that's it.
- dang 5y agoFlamewars are definitely not expected - they're against the rules and something we try to dampen in every way we know. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=by%3Adang%20flame&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- Dylan16807 5y ago
- SavantIdiot 5y agoA bit off topic, but when did a DWORD implicitly become 64bits?
- robertlagrant 5y agoI heard the bird was DWORD.
- veltas 5y agoLots of ISAs consider 32-bit to be a 'word'. And now they have the same problem that Intel already encountered that it's easier to start referring to a new larger native word size as a 'double-word' and the cycle continues...
- SavantIdiot 5y agoClearly I'm stuck in late-1980's-x86-ASM-land. BYTE, WORD, DWORD[, QWORD].
- dragontamer 5y agoHmmm... I think this argument is solid. Albeit biased from GMP's perspective, but bignums are used all the time in RSA / ECC, and probably other common tasks, so maybe its important enough to analyze at this level. 2-instructions to work with 64-bits, maybe 1 more instruction / macro-op for the compare-and-jump back up to a loop, and 1 more instruction for a loop counter of somekind? So we're looking at ~4 instructions for 64-bits on ARM/x86, but ~9-instructions on RISC-V. The loop will be performed in parallel in practice however due to Out-of-order / superscalar execution, so the discussion inside the post (2 instruction on x86 vs 7-instructions on RISC-V) probably is the closest to the truth. ---------- Question: is ~2-clock ticks per 64-bits really the ideal? I don't think so. It seems to me that bignum arithmetic is easily SIMD. Carries are NOT accounted for in x86 AVX or ARM NEON instructions, so x86, ARM, and RISC-V will probably be best. I don't know exactly how to write a bignum addition loop in AVX off the top of my head. But I'd assume it'd be similar to the 7-instructions listed here, except... using 256-bit AVX-registers or 512-bit AVX512 registers. So 7-instructions to perform 512-bits of bignum addition is 73-bits-per-clock cycle, far superior in speed to the 32-bits-per-clock cycle from add + adc (the 64-bit code with implicit condition codes). AVX512 is uncommon, but AVX (256-bit) is common on x86 at least: leading to ~36-bits-per-clock tick. ---------- ARM has SVE, which is ambiguous (sometimes 128-bits, sometimes 512-bits). RISC-V has a bunch of competing vector instructions. .......... Ultimately, I'm not convinced that the add + adc methodology here is best anymore for bignums. With a wide-enough vector, it seems more important to bring forth big 256-bit or 512-bit vector instructions for this use case? EDIT: How many bits is the typical bignum? I think add+adc probably is best for 128, 256, or maybe even 512-bits. But moving up to 1024, 2048, or 4096 bits, SIMD might win out (hard to say without me writing code, but just a hunch). 2048-bit RSA is the common bignum, right? Any other bignums that are commonly used? EDIT2: Now that I think of it, addition isn't the common operation in RSA, but instead multiplication (and division which is based on multiplication).
- serentty 5y ago> RISC-V has a bunch of competing vector instructions. There is only one standard V extension. Alibaba made a chip with a prerelease version of that V extension which is thus incompatible with the final version, but in practice that just means that the vector unit on that chip is not used because it is incompatible, not that there are now competing standards
- pcwalton 5y agoDoesn't RISC-V have an add-with-carry instruction as part of the vector extension? I see it listed here: https://github.com/riscv/riscv-v-spec/releases/tag/v1.0 https://github.com/riscv/riscv-v-spec/releases/tag/v1.0
- monocasa 5y agoAfaict that's only for operations the vector register file. Most of the complaints about the lack of addc/subc are around how they're heavily used in JITs for languages that want to speculatively optimize multi precision arthimetic into the integer register file for their regular integer ops. JavaScript, a lot of Lisps, the MLs all fit into that space.
- pcwalton 5y agoSure, but this email is in the context of GMP, which should be using the vector extension, no?
- monocasa 5y agoI don't think so; most of the users I know of for the integer side of GMP are compilers/runtimes. An apt rdepends on the gmp packages in Ubuntu only shows stuff like ocaml, and I know gcc vendors it. Edit: Another place you see this kind of arthimetic is crypto, but those specific use cases (Diffie Hellman, RSA, a few others) don't tend to be vectorized. You have one op you're trying to work through with large integers, and there's the carry dependency on each partial op. The carry depdent crypto algorithms aren't typically vectorisable.
- pcwalton 5y agoI'm a bit confused. First, JavaScript doesn't use bignum arithmetic much. JS Numbers are doubles. JavaScript JITs do speculatively use integers where possible, but this workload doesn't need add-with-carry, just efficient overflow detection (which is a problem with RISC-V, but not the problem we're discussing here). It could be a problem for Lisp and ML, but these are hardly common languages. And as for crypto, couldn't you use the vector registers to do scalar ops?
- bell-cot 5y agoRather than glib hand-waving in front of the chalkboard...might there be a decent piece or few of RISC V hardware, which could actually be compared to non-RISC V hardware with similar budgets (for design work, transistor count, etc.) - to see how things work out when running substantial pieces of decently-compiled code?
- ksec 5y agoThe unwritten rule of HN: You do not criticise The Rusted Holy Grail and the Riscy Silver Bullet.
- rbanffy 5y agoOr, if you do, you'd better be absolutely right, or people will tear your argument to shreds.
- sophacles 5y agoIt's a sad state - who wants to have their random inaccurate theory debunked by facts? Imagination land is way, way more fun.
- mhh__ 5y agoRust is pretty well regarded technically here, but RISC-V is mostly a "wouldn't it be nice" rather than something most commenters seem to be super knowledgeable about. Many people still think that RISC-V implies an open source implementation, for example.
- DeathArrow 5y agoYou can do worse than that. You can take the Linux's name in vain.
- Shadonototra 5y agoWho changed the title? Moderators where are you?
- marcodiego 5y agoSo, how meaningful is the "projected score of 11+ SPECInt2006/GHz" as claimed here: https://www.sifive.com/press/sifive-raises-risc-v-performance-bar-with-new-best-in-class https://www.sifive.com/press/sifive-raises-risc-v-performanc... ?
- Symmetry 5y agoI expect it to be true but not very meaningful. Clock efficiency in isolation isn't any more useful a figure of merit than clock frequency in isolation. The fact that they're using a metric like that makes me pessimistic about the chip.
- fhood 5y agoOh wow, everybody else is debating the specific intricacies of the design decisions, and I'm here wondering why you would complain about not enough instructions in an architecture with "RISC" in the name.
- adrian_b 5y agoThe RISC idea was to not include in the ISA instructions so complex that they would require a multi-cycle implementation. The minimum duration of the clock cycle of a modern CPU is essentially determined by the duration of a 64-bit integer addition/subtraction, because such operations need a latency of only 1 clock cycle to be useful. Operations that are more complex than 64-bit integer addition/subtraction, e.g. integer multiplications or floating-point operations, need multiple cycles, but they are pipelined so that their throughput remains at 1 per cycle. So 64-bit addition/subtraction is certainly expected to be included in any RISC ISA. The hardware adders used for addition/subtraction provide, at a negligible additional cost, 2 extra bits, carry and overflow, which are needed for operations with large integers and for safe operations with 64-bit integers. The problem is that the RISC-V ISA does not offer access to those 2 bits and generating them in software requires a very large cost in execution time and in lost energy in comparison with generating them in hardware. I do not see any relationship between these bits and the RISC concepts, omitting them does not simplify the hardware, but it makes the software more complex and inefficient.
- jhallenworld 5y agoWhat if the multi-precision code is written in C? You can detect carry of (a+b) in C branch-free with: ((a&b) | ((a|b) & ~(a+b))) >> 31 So 64-bit add in C is: f_low = a_low + b_low c_high = ((a_low & b_low) | ((a_low | b_low) & ~f_low)) >> 31 f_high = a_high + b_high + c_high So for RISC-V in gcc 8.2.0 with -O2 -S -c add a1,a3,a2 or a5,a3,a2 not a7,a1 and a5,a5,a7 and a3,a3,a2 or a5,a5,a3 srli a5,a5,31 add a4,a4,a6 add a4,a4,a5 But for ARM I get (with gcc 9.3.1): add ip, r2, r1 orr r3, r2, r1 and r1, r1, r2 bic r3, r3, ip orr r3, r3, r1 lsr r3, r3, #31 add r2, r2, lr add r2, r2, r3 It's shorter because ARM has bic. Neither one figures out to use carry related instructions. Ah! But! There is a gcc macro: __builtin_uadd_overflow() that replaces the first two C lines above: c_high = __builtin_uadd_overflow(a_low, b_low, &f_low); So with this: RISC-V: add a3,a4,a3 sltu a4,a3,a4 add a5,a5,a2 add a5,a5,a4 ARM: adds r2, r3, r2 movcs r1, #1 movcc r1, #0 add r3, r3, ip add r3, r3, r1 RISC-V is faster.. EDIT: CLANG has one better: __builtin_addc(). f_low = __builtin_addcl(a_low, b_low, 0, &c); f_high = __builtin_addcl(a_high, b_high, c, &junk); x86: addl 8(%rdi), %eax adcl 4(%rdi), %ecx ARM: adds w8, w8, w10 add w9, w11, w9 cinc w9, w9, hs RISC-V: add a1, a4, a5 add a6, a2, a3 sltu a2, a2, a3 add a6, a6, a2
- volta83 5y ago> RISC-V is faster.. I find it funny that you make the same pitfall than the author did. Faster on which CPU? The author doesn't measure on any CPU, so here there are dozens of people hypothesizing whether fusion happens or not, and what the impact is.
- brutal_chaos_ 5y ago> Faster on which CPU? Perhaps faster means fewer instructions in this instance? Considering number of instructions is what has been discussed.
- xondono 5y agoExperimenting with RISC-V is one of those things I keep postponing. For those are more versed, is this really a general problem? I was under the impression that the real bottleneck is memory, and things like this would be fixed in real applications through out of order execution, and that it payed off having simpler instructions because compilers had more freedom to rearrange things.
- fwsgonzo 5y agoRISC-V is completely fine, heavily based on research and well thought out. It does have pros and cons like any other architecture, and for what it does well, it does it really well!
- dlsa 5y agoI noticed high and low in there so those code snippets look like 32 bit code, at least to me. Is that even a fair comparison given the arm and x86 versions used as examples of "better" were 64 bit? If we're really comparing 32 and 64 and complaining that 32 bit uses more instructions than 64, perhaps we should dig out the 4 bit processors and really sharpen the pitchforks. Alternatively, we could simply not. Comparing apples to oranges doesn't really help. From the article: Let's look at some examples of how Risc V underperforms. First, addition of a double-word integer with carry-out: add t0, a4, a6 // add low words sltu t6, t0, a4 // compute carry-out from low add add t1, a5, a7 // add hi words sltu t2, t1, a5 // compute carry-out from high add add t4, t1, t6 // add carry to low result sltu t3, t4, t1 // compute carry out from the carry add add t6, t2, t3 // combine carries Same for 64-bit arm: adds x12, x6, x10 adcs x13, x7, x11 Same for 64-bit x86: add %r8, %rax adc %r9, %rdx
- adrian_b 5y agoThe comparison is completely fair, because on RISC-V there is no better way to generate the carries required for computations with large integers. You cannot generate a carry with a 64-bit addition, because it is lost and you cannot store it. You should take into account that the libgmp authors have a huge amount of experience in implementing operations with large integers on a very large number of CPU architectures, i.e. on all architectures supported by gcc, and for most of those architectures libgmp has been the fastest during many years, or it still is the fastest.
- dlsa 5y agoSo the 32 bit code and the 64 bit code is equally inefficient in your opinion?
- jasonhansel 5y agoOne thing that bothers me: RISC-V seems to use up a lot of the available instruction set space with "HINT" instructions that nobody has (yet) found a use for. Is it anticipated that all of the available HINTs will actually be used, or is the hope that the compressed version of the instruction set will avoid the wasted space?
- oneplane 5y agoAll of the discussions about instruction sets and "mine is better than yours" or "anyone else could do better in a small amount of time" are useless considering those arguments, if true, haven't actually resulted in any free ISA being available broadly, embraced broadly and hardware implementing that ISA being available. It doesn't matter how great something else could be in theory if it doesn't exist or doesn't meet the same scale and mindshare (or adoption).
- kelnos 5y ago> My conclusion is that Risc V is a terrible architecture. Kinda stopped reading here. It's a pretty arrogant hot take. I don't know this guy, maybe he's some sort of ISA expert. But it strains credulity that after all this time and work put into it, RISC-V is a "terrible architecture". My expectation here is that RISC-V requires some inefficient instruction sequences in some corners somewhere (and one of these corners happens to be OP's pet use case), but by and large things are fine. And even then, I don't think that's clear. You're not going to determine performance just by looking at a stream of instructions on modern CPUs. Hell, it's really hard to compare streams of instructions from different ISAs.
- mhh__ 5y agoCalling it terrible is definitely something from the book of Linus T. Bad? Quite possible, it was meant as a teaching ISA initially IIRC, but terrible? Who knows.
- foxfluff 5y agoThat's the difficulty here, people are already arguing past each other because nobody seems to agree what the ISA is for. If you look at the early history of RISC-V, it does indeed look like as something built for teaching. But I don't think that use case warrants all the hype around it. So how did all the hype form, and why is it that there are people seemingly hyping it as the next-gen dream-come-true super elegant open developed-with-hindsight ISA that will eventually displace crufty old x86 and proprietary ARM while offering better performance and better everything? Of course that just baits you into arguing about its potential performance. And don't worry if it doesn't have all the instructions you need for performance yet, we'll just slap it with another extension and it totally won't turn into a clusterfuck with a stench of legacy and numerous attempts at fixing it (coz' remember, hindsight)! And then if you question its potential, you'll get someone else arguing that no no, it's not a high performance ISA for general use in desktops / servers, it's just an extensible ISA that companies can customize for their special sauce microcontrollers or whatever. Of course it's all armchair speculation because there are no high performance real world implementations and there aren't enough experts you can trust.
- socialdemocrat 5y agoRISC V is an opinionated architecture and that is always going to get some people fired up. Any technology that aims for simplicity has to make hard choices and trade offs. It isn’t hard to complain about missing instructions when there are less than 100 of them. Meanwhile nobody will complain about ARM64 missing instructions because it had about 1000 of them. Therein lies the problem. Nobody ever goes out guns blazing complaining about too many instructions despite the fact that complexity has its own downsides. RISC-V has been designed aggressively to have minimal ISA to leave plenty of room to grow, and require minimal number of transistors for a minimal solution. Should this be a showstopper down the road, then there will be plenty of space to add an extensions that fixes this problem. Meanwhile embedded systems paying a premium for transistors are not going to have to pay for these extra instructions as only 47 instructions have to be implemented in a minimal solution.
- klelatti 5y agoFair point but Arm is not just Arm64 - if you want a simple low cost ISA then there is Cortex-M.
- imtringued 5y agoIt reminds me of nutrition advice. The 70s said X is evil or bad. Then we discover, X doesn't matter. I think in 10-20 years everyone will agree that all the "bad" RISC-V decisions don't matter. The same way x86 (CISC) was supposed to be bad because of legacy/backwards compatibility.
- DeathArrow 5y agoI remember Jim Keller saying that for every ISA just 7 or 8 instructions are important and they are used for like 99% of the code. Load, store and the like.
- rep_lodsb 5y agoThat doesn't mean we should eliminate everything else. Many operations that would require 3 or more simple instructions are easy enough to implement in hardware, so that they take no more time than one of these simple instructions. This can be a big win even if they are only used 1% of the time.
- deleted 5y ago[deleted]
- tomxor 5y ago> Let's look at some examples (7 instructions vs 2 vs 2) Isn't this the classic RISC vs CISC problem? Comparing x86/ARM to RISC-V feels like Apples to Grains of Rice. If RISC-V was born out of a need for an open source embedded ISA, would the ISA not need to remain very RISC-like to accommodate implementations with fewer available transistors... Or is this an outdated assumption?
- mhh__ 5y agoIt's not exactly RISC VS CISC. Maybe SISC - "Simplified" instruction set computing, perhaps. ARM isn't exactly super complicated in this particular aspect (it is elsewhere), but in this case the designers basically chose to make branches simpler at the expense of code that needs to check overflows (or flags more generally) RISC-V was born partly out of a desire for a teaching ISA, also, so simplicity is a boon in that context too.
- bArray 5y ago> I believe that an average computer science student could come up with a better instruction set that Risc V in a single term project. When you hear the "<person / group> could make a better <implementation> in <short time period>" - call them out. Do it. The world will not shun a better open license ISA. We even have some pretty awesome FPGA boards these days that would allow you to prototype your own ISA at home. In terms of the market - now is an exceptionally great time to go back to the design room. It's not as if anybody will be manufacturing much during the next year with all of the fab labs unable to make existing chips to meet demand. There is a window of opportunity here. > It is, more-or-less a watered down version of the 30 year old Alpha ISA after all. (Alpha made sense at its time, with the transistor budget available at the time.) As I see it, lower numbers of transistors could also be a good thing. It seems blatantly obvious at this point that multi-core software is not only here to stay, but is the future. Lower numbers of transistors means squeezing more cores onto the same silicon, or implementing larger caches, etc. I also really like the Unix philosophy of doing one simple thing well. Sure, it could have some special instruction that does exactly your use case in one cycle using all the registers, but that's not what has created such advances in general purpose computing. > Sure, it is "clean" but just to make it clean, there was no reason to be naive. I would much rather we build upon a conceptually clean instruction set, rather than trying to hobble together hacks on top of fundamentally flawed designs - even at the cost of performance. It's exactly these hobbled conceptual hacks that have lead to the likes of spectre and meltdown vulnerabilities, when the instruction sets become so complicate that they cannot be easily tested.
- kbelder 5y agoYeah. I don't have a dog in this fight... I don't have strong opinions either way, and this is one of those arguments that will be settled by reality after some time goes by. But the author making an argument like that... > I believe that an average computer science student could come up with a better instruction set that Risc V in a single term project. Pretty much blew their credibility. It's obviously wrong, and a sensible, fair person wouldn't write it.
- imtringued 5y ago
- kazinator 5y agoGodbolt: typedef __int128_t int128_t; int128_t add(int128_t left, int128_t right) { return left + right; } GCC 10, -O2, RISC-V: add(__int128, __int128): mv a5,a0 add a0,a0,a2 sltu a5,a0,a5 add a1,a1,a3 add a1,a5,a1 ret ARM64: add(__int128, __int128): adds x0, x0, x2 adc x1, x1, x3 ret This issue hurts the wider types that are compiler built-ins. Even though C has a programming model that is devoid of any carry flag concept, canned types like a 128 bit integer can take advantage of it. Portable C code to simulate a 128 bit integer will probably emit bad code across the board. The code will explicitly calculate the carry as an additional operand and pull it into the result. The RISC-V won't look any worse, then, in all likelihood. (The above RISC-V instruction set sequence is shorter than the mailing list post author's 7 line sequence because it doesn't calculate a carry out: the result is truncated. You'd need a carry out to continue a wider addition.)
- akimball 5y agoThe sad thing is that many people will just read the headline or the original email and walk away misinformed and indeed disinformed
- adapteva 5y agoOver 200 comments and not a single benchmark comparison, if only there was some way to settle this argument...sigh.
- deleted 5y ago[deleted]
- DeathArrow 5y agoTechnical lead for SoC architecture at Nokia dismissed Risc V: https://www.quora.com/Is-RISC-V-the-future/answer/Heikki-Kultala-1 https://www.quora.com/Is-RISC-V-the-future/answer/Heikki-Kul...
- throwaway4good 5y agoInside AI you have all these standard data sets (ImageNet, MNIST etc.) that act as a benchmark for how well an algorithm performs within a given area (character recognintion, image recognintion etc.). Perhaps something similar is needed within ISAs / CPUs ? Say an OS kernel, a ZIP-algorithm, Mandelbrot, Fizz-buzz ... could measure code compactness but also performance and energy usage.
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- usr1106 5y agoThe code given is arbitrary precision addition. How often do you need that in general computing? Hardly often enough to make a measurable difference. Whether the similar awkwardness applies to a lot of other code or not is not being told by this isolated case.
- mda 5y agoHonestly in my eyes, author loses all credibility after saying this: "I believe that an average computer science student could come up with a better instruction set that Risc V in a single term project" Utter horse manure.
- YesThatTom2 5y agoI call this "benchmark by visual inspection". It is completely useless. Yet, many top devs that I know seem to think that they can emulate a complex chip in their head better than... the chip itself.
- rep_lodsb 5y agoCarry flag and overflow checking? We don't need those things because C doesn't support it! That sadly seems to be the kind of thought process behind RISC-V, and a lot of other "modern" computing: Everything should be written in C, or some scripting language implemented in C. Writing safe code is easy, just wrap everything in layers of macros that the compiler will magically optimize away, and if it doesn't, computers are fast enough anyway, right? The mark of a real programmer is that every one of their source files includes megabytes of headers defining things like __GNU__EXTENSION_FOO_BAR_F__UNDERSCORE_. You say your processor has a single instruction to do some extremely common operation, and want to use it? You shouldn't even be reading a processor manual unless you are working on one of the two approved compilers, preferably GCC! If you are very lucky, those compiler people that are so much smarter than you could hope to be, have already implemented some clever transformation that recognizes the specific kind of expression produced by a set of deeply nested macros, and turns them into that single instruction. In the process, it will helpfully remove null pointer checks because you are relying on undefined behaviour somewhere else. You say you'll do it in assembly? For Kernighan's sake, think about portability!!! I mean, portable to any other system that more or less looks the same as UNIX, with a generous sprinkling of #ifdefs and a configure script that takes minutes to run. Implement a better language? Sure, as long as the compiler is written in C, preferably outputs C source code (that is then run through GCC), and the output binary must of course link against the system's C library. You can't do it any other way, and every proper UNIX - BSD or Mac OS X - will make it literally impossible by preventing syscalls from any other piece of code. IMO this is like a cultural virus that seems to have infected everything IT-related, and I don't exactly understand why. Sure, having all these layers of cruft down below lets us build the next web app faster, but isn't it normal to want to fix things? Do some people actually get a sense of satisfaction out of saying "It is a solved problem, don't reinvent the wheel"? Or do they want to think that their knowledge of UNIX and C intricacies is somehow the most important, fundamental thing in computer science?