3 ms·
Unrelated, but somehow similar: I really hate it that it's not possible to force gcc to transform things like this into a conditional move: x > c ? y : 0.; It
by toredo1729_2 2y ago
Unrelated, but somehow similar: I really hate it that it's not possible to force gcc to transform things like this into a conditional move:
x > c ? y : 0.;
It annoyed me many times and it still does.
- fweimer 2y agoWhat do you mean? Do you want to annotate the condition as unpredictable, so that the compiler always assumes that a conditional move is beneficial? (Compilers obviously do this transformation, including GCC, but it is not always beneficial, especially on x86-64.)
- IshKebab 2y agoAnd it's not always possible! E.g. most RISC-V CPUs don't support it yet.
- dzaima 2y agoEh, it takes ~3-4 instrs to do a branchless "x ? y : z" on baseline rv64i (depending on the format you have the condition in) via "y^((y^z)&x)", and with Zicond that only goes down to 3 instrs (they really don't want to standardize GPR instrs with 3 operands so what Zicond adds is "x ? y : 0" and "x ? 0 : y" ¯\_(ツ)_/¯; might bring the latency down by an instr or two though).
- IshKebab 2y agoIt's more about removing branches than instruction counts or latency.
- dzaima 2y agoThe "y^((y^z)&x)" method is already branchless, and close in performance to the Zicond variant, is my point; i.e. Zicond doesn't actually add much.
- IshKebab 2y agoAre you sure? As soon as you add actual computations in you're heading through the whole execution pipeline & forwarding network, tying up ALUs, etc. Zicond can probably be handled without all that. Also that isn't actually equivalent since `x` needs to be all 1s or all 0s surely? Neither GCC nor Clang use that method, but they do use Zicond.
- dzaima 2y agoZicond's czero.eqz & czero.nez (& the `or` to merge those together for the 3-instr impl of the general `x?y:z`) still have to go through the execution pipeline, forwarding network, an ALU, etc just as much as an xor or and need to. It's just that there's a shorter dependency chain and maybe one less instr. Indeed you may need to negate `x` if you have only the LSB set in it; hence "3-4 instrs ... depending on the format you have the condition in" in my original message. I assume gcc & clang just haven't bothered considering the branchless baseline impl, rather than it being particularly bad. Note that there's another way some RISC-V hardware supports doing branchless conditional stores - a jump over a move instr (or in some cases, even some arithmetic instructions), which they internally convert to a branchless update.
- toredo1729_2 2y agoYes, that would be great. It's not always benefical, but in some (rare, but for me important) cases it's better. Currently, the only way to ensure a conditional move is used, is to use inline assembly. This is not portable and also less maintainable than a "proper" solution.
- tjalfi 2y agoclang has the __builtin_unpredictable() intrinsic[0] for this purpose. [0] https://clang.llvm.org/docs/LanguageExtensions.html#builtin-unpredictable https://clang.llvm.org/docs/LanguageExtensions.html#builtin-...
- flohofwoe 2y agoSeems to work just fine on gcc and clang? https://www.godbolt.org/z/ffEvvjhz8 https://www.godbolt.org/z/ffEvvjhz8 PS: and it also doesn't matter whether a ternary is used or a traditional if (as one would expect): https://www.godbolt.org/z/zjb4KdqvK https://www.godbolt.org/z/zjb4KdqvK (the float version also appears to not use branches: https://www.godbolt.org/z/98bdheKK4 https://www.godbolt.org/z/98bdheKK4) For such simple expression I would expect the compiler to pick the right output pattern based on the target CPU though...
- dzaima 2y agoNot always - https://www.godbolt.org/z/zYxeahf3T https://www.godbolt.org/z/zYxeahf3T. And for any modern (as in, made in the last two decades) x86 processor the branchless version will be hilariously better if the condition is unpredictable (which is a thing the compiler can't know by itself, hence wanting to have an explicit way to request a conditional move instr) and the per-branch code takes less than like multiple dozens of cycles.
- dzaima 2y agoWorse, doing one of the idioms for a conditional move ends up getting gcc to actually produce a conditional move, but clang doesn't, even with its __builtin_unpredictable: https://www.godbolt.org/z/bq9axzvjG https://www.godbolt.org/z/bq9axzvjG
- ryao 2y agoYou want to pass -mllvm -x86-cmov-converter=false. I assume LLVM has a pass to undo conditional moves on x86 whenever a heuristic determines skipping a calculation by branching is cheaper than doing the calculation and using a conditional move. Unfortunately, the heuristic that calculates the expense often gets things wrong. That is why OpenZFS passes -mllvm -x86-cmov-converter=false to Clang for certain files where the LLVM heuristic was found to do the wrong thing: https://github.com/openzfs/zfs/commit/677c6f8457943fe5b56d7aa8807010a104563e4a https://github.com/openzfs/zfs/commit/677c6f8457943fe5b56d7a... There is an open LLVM issue regarding this: https://github.com/llvm/llvm-project/issues/62790 https://github.com/llvm/llvm-project/issues/62790 The issue explains why __builtin_unpredictable() does not address the problem. In short, the metadata is dropped when an intermediate representation is generated inside clang since the IR does not have a way to preserve the information.