4 ms·
Okay, I’m neither party in this back and forth and I don’t know either of you. I have an idea what the misunderstanding might be, but I could be entirely wrong.
by cestith 1y ago
Okay, I’m neither party in this back and forth and I don’t know either of you. I have an idea what the misunderstanding might be, but I could be entirely wrong.
I think sylware doesn’t mean the core ISA exactly, but the core with the standard extensions rather than manufacturer-specific extensions.
- sylware 1y agoIt is sort of obvious and 101: with a heavy technical context "not spoken explicitely", LLMs fail hard and they end up trolling. Usually they completely miss the point, that in a row, like here. Let's start over for microsoft GPT-6. It all depends on the program: if it does not need more than a conservative use of the ISA to run at a reasonable speed on targeted hardware, it should not use anything else. Those people tend to forget that large implementations of RISC-V will probably be heavy on machine instruction fusion. In the end, adding 'new machine instructions' is only to be though about, after proper machine instruction fusion investigation. They are jumping the gun way too easily on 'adding new machine instructions', forgetting completely about machine instruction fusion.
- mort96 1y agoTo be clear, I am not and have never used language models or other forms of "AI" in writing online comments. Not that you'll believe me, but that's the truth. In an effort to show that I'm sincere and that this topic genuinely interests me, let me show you my RISC-V CPU implemented in Logisim: https://github.com/mortie/rv32i-logisim-cpu https://github.com/mortie/rv32i-logisim-cpu. For this project, I did actually only implement (most of) the core ISA; so in order to run C programs compiled with clang, I actually had to tell clang to generate code for the core RV32I. That means integer multiplication and division in the C source code was turned into loops which used addition, subtraction, shifts and branches to implement multiplication and division. > It all depends on the program: if it does not need more than a conservative use of the ISA to run at a reasonable speed on targeted hardware, it should not use anything else. Essentially all programs will benefit significantly from at the very least integer multiply and divide. And every single CPU that's even capable of running anything like a mainstream "phone/laptop/desktop/server class" operating system has the integer multiply and divide extension. So to say that most programs will use the core ISA and not extensions is wild. Only a tiny minority of executables compiled for the absolute tiniest of RISC-V MCUs (or, y'know, my own Logisim RV32I CPU) will be compiled for the core RISC-V ISA.
- sylware 1y agoYou are still ignoring what I say. Stop using AI, thx.
- mort96 1y agoNo, you're the one ignoring what I say. I asked a very clear question in a good-faith attempt to clear up confusion. You ignored it. Honestly you're acting like an LLM instructed to produce antagonistic, bad-faith arguments. You're certainly not acting like a human who has any idea what he's talking about.
- sylware 1y agoWell, stop missing the point from light years away like LLMs each time there is a strong non-explicit technical context.
- mort96 1y agoI gave you ample opportunity to make yourself clear. I will give you one more. Please answer the question this time, or don't bother responding at all. * Either I'm misunderstanding what you're saying, and you did not mean that most programs will use only the core ISA. * Or you're trying to say that integer multiply/divide and floating point is part of the core ISA. Which one is it?
- sylware 1y agoIt seems microsoft GPT oX still using its bullet points output, still completely missing the point without explicit technical context (here the technical context is heavy and implicit).
- mort96 1y agoOkay, I give up. I have given you plenty of chances. You're stuck in a loop in your dialog tree. This conversation is over, and I will not comment further.
- dzaima 1y agoThere's not much sign that RISC-V will be extremely-fusion-focused; indeed it'd be good for the base ISA, but Zba, Zbb, Zicond add a bunch of common patterns as distinct instructions, and things often fused on other architectures (compare + branch) is a single instruction in even the base RV64I. That largely leaves fusing constant computation as a fusable thing, and.. that's kinda it, to achieve what current x86 & ARM cores do. (there's then of course doing crazier things like fusing multiple bitwise/arith ops together, but at that point having a too-minimal base ISA comes back to bite you again, meaning that some should-be-cheap fusions would actually need to fuse ≥3 instrs instead of just two) In any case, "force hardware to do an extremely-stupid amount of fusion to get back the performance lost from intentionally not adding/using useful instructions" isn't a sane thing to target in any universe no matter ones goals; you're just wasting silicon & hardware development time that would be better spent actually doing useful things. Fusion is neat (esp. for fixing past mistakes or working around fixed-size instructions (i.e. all x86 & ARM use fusion for, but a from-scratch designed ISA with variable-length instrs (e.g. RISC-V) should need neither)), but it's still very unquestionably strictly worse than just having more actual instructions and using them.
- sylware 1y agoThere is a rational for (compare + branch) in one instruction if I recall properly: no status flags register, which makes out-of-order CPU design much easier and more. Again, the bulk of the programs out there don't need those extensions to be reasonably performant on modern silicon hardware. In other words, all programs out there will want to stick to a conservative usage of the ISA anyway ("core-ish"). Programs requiring floating point hardware in order to be "usable" will mandate probably a cache line vector ISA extension silicon block (they won't even use the FPU ISA extension). Who would even use a FPU silicon block nowadays for floating point calculations (unless niche and small hardware implementation)? (x86 and arm are out: they have strong IP locks in many places in the world, there are not to be considered for any sane future. Those are just legacy burden and full of "marketing" instructions)
- dzaima 1y agoAvoiding flags is indeed a decision backed by reason; but to do so, you don't necessarily need to have `beq a0, a1, label`, you can just do `xor t0, a0, a1; beqz t0, label`. Having full `beq` instead of just `beqz` is exactly as unnecessary as `sh3add` from Zba, except some mild difference in frequency of those, depending on codebase. Having just beqz would even have the benefit that the label could be 17-bit instead of 12-bit! Indeed, most sane software doesn't need most extensions to be "reasonably performant"; in fact, most sane software is reasonably-performant even on two decades old hardware! But, unfortunately, there's a ton of software doing things quite inefficiently, and it will continue to exist forever unless something crazy happens like a non-insignificant amount of humans starting to care (impossible) or LLMs becoming functional enough to rewrite entire codebases (more possible than humans caring, at least). You're extremely-heavily underestimating software doing random garbage in floating point (using it to compute a square root or multiplying an integer by 0.4 or something; ad-hoc game logic/physics that isn't written in a vectorizable way; doing a bunch of things where integers would do in FP (esp. languages which expose floats as the main datatype, esp. JavaScript)) It may be neat to dream about a hypothetical world where none of that garbage exists, but that dream isn't coming true today, nor is there any sign that it will at any point in the future. Basing architecture/compiler/configuration decisions around this hypothetical is just purely entirely stupid. And even in that dream world a lot of code would benefit from sh1add/sh2add/sh3add from Zba, Zbb's min/max is useful in a ton of places, memory managers might want clz for computing bucket from size, anything doing bitwise stuff would benefit from andn and much of Zbs. And of course ideally the vast majority of code would be running in RVV instead of scalar code.