5 ms·
This isn't true, this form of conditionals can be compiled into cmov type of instructions, which is faster than regular jump if condition.
by enedil 5y ago
This isn't true, this form of conditionals can be compiled into cmov type of instructions, which is faster than regular jump if condition.
- ncann 5y agoIf the if/else is simple the compiler should be able to optimize that anyway.
- dataflow 5y ago> This isn't true, this form of conditionals can be compiled into cmov type of instructions, which is faster than regular jump if condition. IIRC cmov is actually quite slow. It's just faster than an unpredictable branch. Most branches have predictability so you generally don't want a cmov. Speaking of which, a couple questions regarding this for anyone who might know: 1. Can you disable cmov on x64 on any compiler? How? 2. Why is cmov so slow? Does it kill register renaming or something like that?
- dgrunwald 5y agocmov itself isn't slow, it has a latency of 2 cycles on Intel; and only 1 cycle on AMD (same speed as an add). However, cmov has to wait until all three inputs (condition flag, old value of target register, value of source register) are available, even though one of those inputs ends up going unused. A correctly predicted branch allows the subsequent computation (using of the result of the ?: operator) to start speculatively after waiting only for the relevant input value, without having to wait for the condition or the value on the unused branch. This could sometimes save hundreds of cycles if an unused input is slow due to a cache miss.
- bigiain 5y agoI wonder if there's anyone on earth who needs nicely formatted human readable file sizes that's worried about the difference between one or two cpu cycle branching instructions? There might be a few guys at FAANG who have a planet-scale use case for human readable file sizes. But surely "performance optimising" this is _purely_ code golf geekiness? (Which is a perfectly valid reason to do it, but I'm gonna choose the most obvious to the next progerammer reading it version over one that 50% or 500% or 5000% fast in almost any use case I can think I'm like to need this... I mean, it's only looking for 6 prefixes "KMGTPE" a six line case statement would work for most people?)
- bigiain 5y agoActually, I just realised. This is (probably a small part of) why "calculate all sizes" in Mac finder windows is so slow. I already mentioned Apple in FAANG, but I guess someone at Microsoft and people who work on Linux file brokers care too. And whoever maintains the -h flag codepaths in all the Unix-like utils that support it?
- dataflow 5y agoConfused what this has to do with calculating file sizes. Time spent computing file sizes is dwarfed by I/O, right?
- dataflow 5y agoAhh, thank you! Makes sense.
- initplus 5y agoCMOV is slow because x86 processors will not speculate past a CMOV instruction. They do speculate past conditional jumps, so those are more performant. This same property makes CMOV useful in Spectre mitigation, see https://llvm.org/docs/SpeculativeLoadHardening.html https://llvm.org/docs/SpeculativeLoadHardening.html Keeping CMOV slow is now an important security feature.
- dataflow 5y agoThey don't speculate past a CMOV at all? Like even if the next instruction has nothing to do with the CMOV's output?
- saghm 5y agoI think out of order processing is considered different than speculative execution, but I could be remembering my architecture class wrong
- colejohnson66 5y agoOut-of-order just means it can rearrange the decoded uops in a way to keep the execution units at full capacity. So, if an instruction needs the ALU, but it’s busy, and the next one needs the AGU (address generation unit) and doesn’t depend on the results of the ALU one, it can “dispatch” the AGU one while the ALU one waits for the pipeline to move. Speculative execution refers more towards the decoder/uop generation side of the processor (the “in-order” side). A normal “in-order” processor, upon encountering a conditional jump, would wait until the pipeline is finished to check if it should jump or not. It does it by inserting “bubbles” into the pipeline - essentially doing nothing but waiting. Speculative execution (or branch prediction) would say, “I think the branch will be taken based on X, Y, Z,” and then keep the pipeline full in the process. If the prediction was right, congratulations! You just saved dozens of clock cycles that otherwise would’ve been wasted. If it was wrong, no worries. The pipeline is then flushed; all the speculated instructions’ results are tossed (before they’re “written back”). Then the processor resumes operation on the correct branch. Speculative execution doesn’t necessitate an out-of-order architecture, and visa-versa. Just a pipelined one. It’s perfectly possible to have an out-of-order architecture that doesn’t speculate, or a speculative one that is completely “in-order”, but they work hand-in-hand, and it makes sense to have both if you have one.
- rot13xor 5y agoThis email thread from Linus might be interesting: https://yarchive.net/comp/linux/cmov.html https://yarchive.net/comp/linux/cmov.html
- dataflow 5y agoI think that thread is where I first learned this actually. Didn't remember it until you linked it now, thanks for posting it!
- colejohnson66 5y agoMy understanding of out-of-order (and pipelined) CPUs is limited, but it’s interesting that CMOV isn’t interpreted as a “Jcc over MOV” by the decoder. That would allow using the branch predictor. Would it be too complex or does the microarchitecture not even allow it?
- hvdijk 5y agoBoth ?: and if-else have cases where they can be compiled into cmov type instructions and where they cannot. Given int max(int a, int b) { if (a > b) return a; else return b; }, a decent compiler for X86 will avoid conditional branches even though ?: wasn't used. Given int f(int x) { return x ? g() : h(); }, avoiding conditional branches is more costly than just using them, even though ?: was used.