15 ms·
Don't "optimize" conditional moves in shaders with mix()+step()
- ttoinou 2y agoThanks Inigo ! The second wrong thing with the supposedly optimizer version is that it actually runs much slower than the original version. The reason is that the step() function is actually implemented like this: float step( float x, float y ) { return x < y ? 1.0 : 0.0; } How are we supposed to know what OpenGL functions are emulated rather than calling GPU primitives ?
- Const-me 2y agoThe only way is do what OP did – compile your shader, disassemble, and read the assembly. I do that quite often with my HLSL shaders, learned a lot about that virtual instruction set. For example, it’s interesting GPUs have instruction sincos, but inverse trigonometry is emulated while compiling.
- Waterluvian 2y agoThis is a great question that I see everywhere in programming and I think it is core to why you measure first when optimizing. You generally shouldn’t know or care how a built in is implemented. If do care, you’re probably thinking about optimization. At that point the answer is “measure and find out what works better.”
- TeMPOraL 2y agoEDIT: I see my source of confusion must be that "branch" must have a well-understood hardware-specific meaning that goes beyond the meaning I grew up with, which is that a conditional is a branch, because the path control takes (at the machine code level) is chosen at runtime. This makes a conditional jump a branch by definition. > How are we supposed to know what OpenGL functions are emulated rather than calling GPU primitives? To me the problem was obvious, but then again I'm having trouble with both your and author's statements about it. The problem I saw was, obviously by going for a step() function, people aren't turning logic into arithmetic, they're just hiding logic in a library function call. Just because step() is a built-in or something you'd find used in mathematical paper doesn't mean anything; the definition of step() in mathematics is literally a conditional too. Now, the way to optimize it properly to have no conditionals, is you have to take a continuous function that resembles your desired outcome (which in the problem in question isn't step() but the thing it was used for!), and tune its parameters to get as close as it can to your target. I.e. typically you'd pick some polynomial and run the standard iterative approximation on it. Then you'd just have an f(x) that has no branching, just a bunch of extra additions and multiplications and some "weirdly specific" constants. Where I don't get the author is in insisting that conditional move isn't "branching". I don't see how that would be except in some special cases, where lack of branching is well-known but very special implementation detail - like where the author says: > also note that the abs() call does not become a GPU instruction and instead becomes an instruction modifier, which is free. That's because we standardized on two's complement representation for ints, which has the convenient quality of isolating sign as the most significant bit, and for floats the representation (IEEE-754) was just straight up designed to achieve the same. So in both cases, abs() boils down to unconditionally setting the most significant bit to 0 - or, equivalently, masking it off for the instruction that's reading it. step() isn't like that, nor any other arbitrary ternary operation construct, and nor is - as far as I know - a conditional move instruction. As for where I don't get 'ttoinou: > How are we supposed to know what OpenGL functions are emulated rather than calling GPU primitives The basics like abs() and sqrt() and basic trigonometry are standard knowledge, the rest... does it even matter? step() obviously has to branch somewhere; whether you do it yourself, let a library do it, or let the hardware do it, shouldn't change the fundamental nature.
- ttoinou 2y agoIt kinda does when you’re wondering what’s going on in backstage and working with shaders on multiple OS, drivers and hardware. Now, the way to optimize it properly to have no conditionals, is you have to take a continuous I suspect that we shaders authors really like Clean Math and that’s also why we like to think such “optimizations” with the step function is a nice modification :-)
- mymoomin 2y agoA "branch" here is a conditional jump. This has the issues the article mentions, which branchless programming avoids: Again, there is no branching - the instruction pointer isn't manipulated, there's no branch prediction involved, no instruction cache to invalidate, no nothing. This has nothing to do with whether the behaviour of some instruction depends on its arguments. Looking at the Microsoft compiler output from the article, the iadd (signed add) instruction will get different results depending on its arguments, and the movc (conditional move) will store different values depending on its arguments, but after each the instruction pointer will just move onto the next instruction, so there are no branches.
- burch45 2y agoBranching is different instruction paths, so it requires reading the instructions from different memory that causes a delay jumping to those new instructions rather than plowing ahead on the current stream of instructions. So a conditional jump is a branch but a conditional move is just an instruction that moves one of two values into a register but doesn’t affect what code is executed next.
- dahart 2y ago> the meaning I grew up with, which is that a conditional is a branch A conditional jump is a branch. But a branch has always had a different meaning than a generic “conditional”. There are conditional instructions that don’t jump, e.g. CMP, and the distinction is very important. Branch or conditional jump means the PC can be set to something other than ‘next instruction’. A conditional, such a conditional select or conditional move, one that doesn’t change the PC, is not a branch. > take a continuous function […] Then you’d just have an f(x) that has no branching One can easily implement conditional functions without branching. You can use a compare instruction followed by a Heaviside function on the result, evaluate both sides of the result, and sum it up with a 2D dot product (against the compare result and its negation). That is occasionally (but certainly not always) faster on a GPU than using if/else, but only if the compiler is otherwise going to produce real branch instructions.
- SideQuark 2y agoI’ve never seen a GPU with special primitives for any functions than you’d see in pc style assembly. Every time I’ve looked at a decompiled shader, it’s always been pretty much what you think of in C. Aldo specs like OpenGL specify many intrinsic behavior, which is then implemented as the spec, using standard assembly instructions. Find an online site that decompiles to various architectures.
- account42 2y agoWhy are you supposed to know? Because you care about performance? step being implemented as a libray function on top of a conditional doesn't really say anything about its performance vs being a dedicated instruction. Don't worry about the implementation. Because you are curious about GPU architectures? Look at disassembly, (open source) driver code (including LLVM) and/or ISA documentation.
- doctorhandshake 2y agoI don’t know enough about these implementations to know if this can be interpreted as a blanket ‘conditionals are fine’ or, rather, ‘ternary operations which select between two themselves non-branching expressions are fine’. Like does this apply if one of the two branches of a conditional is computationally much more expensive? My (very shallow) understanding was that having, eg, a return statement on one branch and a bunch of work on the other would hamstring the GPU’s ability to optimize execution.
- TinkersW 2y agoA real branch is useful if you can realistically skip a bunch of work, but this requires all the lanes to agree, on a GPU that means 32 to 64 lanes need to all agree, also for something basic like a few arithmetic ops there is no point.
- dahart 2y agoA GPU/SIMT branch works by running both sides, unless all threads in the thread group (warp/wavefront) make the same branch decision. As long as both paths have at least one thread, the GPU will run both paths sequentially and simply set the active mask of threads for each side of the branch. In other words, the threads that don’t take a given branch sit idle while the active threads do their work. (Note “sit idle” might involve doing all the work and throwing away the result.) If you have two branches, and one is trivial while the other is expensive, and if the compiler doesn’t optimize away the branch already, it may be better for performance to write the code to take both branches unconditionally, and use a conditional assignment at the end. It’s worth knowing that often there are clever techniques to completely avoid branching. Sometimes these techniques are simple, and sometimes they’re invasive and difficult to implement. It’s easy (for me, anyway) to get stuck thinking in a single-threaded CPU way and not see how to avoid branching until you’ve bumped into and seen some of the ways smart people solve these problems.
- toredo1729_2 2y agoUnrelated, but somehow similar: I really hate it that it's not possible to force gcc to transform things like this into a conditional move: x > c ? y : 0.; It annoyed me many times and it still does.
- fweimer 2y agoWhat do you mean? Do you want to annotate the condition as unpredictable, so that the compiler always assumes that a conditional move is beneficial? (Compilers obviously do this transformation, including GCC, but it is not always beneficial, especially on x86-64.)
- IshKebab 2y agoAnd it's not always possible! E.g. most RISC-V CPUs don't support it yet.
- dzaima 2y agoEh, it takes ~3-4 instrs to do a branchless "x ? y : z" on baseline rv64i (depending on the format you have the condition in) via "y^((y^z)&x)", and with Zicond that only goes down to 3 instrs (they really don't want to standardize GPR instrs with 3 operands so what Zicond adds is "x ? y : 0" and "x ? 0 : y" ¯\_(ツ)_/¯; might bring the latency down by an instr or two though).
- ryao 2y agoDo shader compilers have optimization passes to undo this mistake and if not, could they be added?
- DRAGONERO 2y agoI’d expect most vendors do, at least in their closed source drivers. You could also check in the mesa project if this is implemented but it’s definitely possible to do
- ryao 2y agoShader compilers tend to be very latency sensitive, so “it takes too long to run” would be a valid reason why it is not done if it is not done.
- DRAGONERO 2y agoShader compilers mostly use LLVM even though runtime is a constraint, if the pattern is common enough it’s definitely easy to match (it’s just two intrinsics after all) meaning you can do it for cheap in instcombine which you’re going to be running anyway
- ryao 2y agoFor some reason, I feel like this is harder to implement than you expect. The way to find out would be to get a bunch of examples of people doing this “optimizations in shader code, look at the IR generated compared to the optimal version and figure out a set of rules to detect the bad versions and transform it into a good versions. Keep in mind that in the example, the addition operators could be replaced with logical OR operators, so there are definitely multiple variations that need to be detected and corrected.
- DRAGONERO 2y agoI've checked and on "certain vendors" the mix + step is actually (slightly) better: same temp usage, lower instructions/cycles.
- mirsadm 2y agoI've been caught by this. Even Claude/ChatGPT will suggest it as an optimisation. Every time I've measured a performance drop doing this. Sometimes significant.
- WJW 2y agoIs that weird? LLMs will just repeat what is in their training corpus. If most of the internet is recommending something wrong (like this conditional move "optimization") then that is what they will recommend too.
- xbar 2y agoNot weird but important to note.
- diath 2y ago> Even Claude/ChatGPT will suggest it as an optimisation. LLMs just repeat what people on the internet say, and people are often wrong.
- londons_explore 2y agoSo why isn't the compiler smart enough to see that the 'optimised' version is the same? Surely it understands "step()" and can optimize the "step()=0.0" and "step()==1.0" cases separately? This is presumably always worth it, because you would at least remove one multiplication (usually turning it into a conditional load/store/something else)
- NohatCoder 2y agoIt may very well be, it is the type of optimisation where it is quite possible that some compilers may do it some of the time, but it is definitely also possible to write a version that the compiler can't grok.
- Cieric 2y agoThe other part of the optimization issue is that you can't take to long to try anything and everything. Most of the optimizations happen on the driver side, and anything that takes to long will show up as shader compilation stutter. I can't say currently if this is or isn't done, it's just always something you have to think about.
- flowzai4 2y ago[dead]
- magicalhippo 2y agoProcessors change, compilers change. If you care about such details, best to ship multiple variants and pick the fastest one at runtime. As I've mentioned here several times before, I've made code significantly faster by removing the hand-rolled assembly and replacing it with plain C or similar. While the assembly might have been faster a decade or two ago, things have changed...
- dist-epoch 2y agoFunnily enough, this is sort of what the NVIDIA drivers do: they intercept game shaders and replace them by custom ones optimized by NVIDIA. Which is why you see stuff like this in NVIDIA drivers changelog: "optimized game X, runs 40% faster"
- quuxplusone 2y agoI'm sure TFA's conclusion is right; but its argument would be strengthened by providing the codegen for both versions, instead of just the better version. Quote: "The second wrong thing with the supposedly optimizer [sic] version is that it actually runs much slower than the original version [...] wasting two multiplications and one or two additions. [...] But don't take my word for it, let's look at the generated machine code for the relevant part of the shader" —then proceeds to show only one codegen: the one containing no multiplications or additions. That proves the good version is fine; it doesn't yet prove the bad version is worse.
- azeemba 2y agoThe main point is that the conditional didn't actually introduce a branch. Showing the other generated version would only show that it's longer. It is not expected to have a branch either. So I don't think it would have added much value
- idunnoman1222 2y agoUnless you’re writing an essay on why you’re right…
- chrisjj 2y ago> Unless you’re writing an essay on why you’re right… He's writing an essay on why they are wrong. "But here's the problem - when seeing code like this, somebody somewhere will invariably propose the following "optimization", which replaces what they believe (erroneously) are "conditional branches" by arithmetical operations." Hence his branchless codegen samples are sufficient. Further, regarding.the side-issue "The second wrong thing with the supposedly optimizer [sic] version is that it actually runs much slower", no amount of codegen is going to show lower /speed/.
- ncruces 2y agoThe other either optimizes the same, or has an additional multiplication, and it's definitely less readable.
- TinkersW 2y agoIt is weird how long misinformation like this sticks around, the conditional move/select approach has been superior for decades on both CPU & GPU, but somehow some people still write the other approach as an "optimization".
- Lockal 2y ago"Conditional move is superior on CPU" is an oversimplification, in reality it was explained in 2007 by Linus[1] and nothing changed since then. Or in fact, branch predictors are constantly improving[2], while cmov data dependency problem can't be solved. Yes, cmov is better in unpredictable branches, but the universal truth is "programmers are notoriously bad at predicting how their programs actually perform"[3] [1] https://yarchive.net/comp/linux/cmov.html https://yarchive.net/comp/linux/cmov.html [2] https://chipsandcheese.com/p/zen-5s-2-ahead-branch-predictor-unit-how-30-year-old-idea-allows-for-new-tricks https://chipsandcheese.com/p/zen-5s-2-ahead-branch-predictor... [3] https://gcc.gnu.org/onlinedocs/gcc/Other-Builtins.html#index-_005f_005fbuiltin_005fexpect https://gcc.gnu.org/onlinedocs/gcc/Other-Builtins.html#index...
- account42 2y ago> cmov data dependency problem can't be solved Can't is a pretty strong word.
- TinkersW 2y agoI really only care about performance in a SIMD sense, as not using SIMD means you aren't targeting performance to begin with(GPU == SIMD, CPU SSE/AVX == SIMD) Branches fall off vs cmov/select as your lane count increases, they are still useful but only when you are certain the probability all lanes agreeing is reasonably high.
- mahkoh 2y agoSo, if you ever see somebody proposing this float a = mix( b, c, step( y, x ) ); The author seems unaware of float a = mix( b, c, y > x ); which encodes the desired behavior and also works for vectors: The variants of mix where a is genBType select which vector each returned component comes from. For a component of a that is false, the corresponding component of x is returned. For a component of a that is true, the corresponding component of y is returned.
- Thorrez 2y agoThe author doesn't seem to say that mix should be avoided. Just that you shouldn't replace a ternary with step+mix. In your quote, you left out the 2nd half of the sentence: "as an optimization to [ternary]".
- mahkoh 2y agoThe author frames his post to be about education: please correct them for me. The misinformation has been around for 20 years But his education will fail as soon as you're operating on more than scalars. It might in fact do more harm than good since it leads the uneducated to believe that mix is not the right tool to choose between two values.
- dahart 2y agoIf you only pass a boolean 0 or 1 for the “a” mix parameter, when is using mix better than a ternary? Can you give an example? I’m not sure mix is ever the right tool to choose between two values. It’s a great tool for blending two values, for linear interpolation when “a” is between 0 and 1. But if “a” is only 0 or 1, I don’t think mix will help you, and it could potentially hurt if the two values you mix are expensive function calls.
- mahkoh 2y agoa can be a vector of booleans.
- alkonaut 2y agoI wish there was a good way of knowing when an if forces an actual branch rather than when it doesn't. The reason people do potentially more expensive mix/lerps is because while it might cost a tiny overhead, they are scared of making it a branch. I do like that the most obvious v = x > y ? a : b; actually works, but it's also concerning that we have syntax where an if is some times a branch and some times not. In a context where you really can't branch, you'd almost like branch-if and non-branching-if to be different keywords. The non-branching one would fail compilation if the compiler couldn't do it without branching. The branching one would warn if it could be done with branching.
- deleted 2y ago[deleted]
- ajross 2y ago> it's also concerning that we have syntax where an if is some times a branch and some times not. That's true on scalar CPUs too though. The CMOV instruction arrived with the P6 core in 1995, for example. Branches are expensive everywhere, even in scalar architectures, and compilers do their best to figure out when they should use an alternative strategy. And sometimes get it wrong, but not very often.
- masklinn 2y agoFor scalar CPUs, historically CMOV used to be relatively slow on x86, and notably for reliable branching patterns (>75% reliable) branches could be a lot faster. cmov also has dependencies on all three inputs, so if there's a high level of bias towards the unlikely input having a much higher latency than the likely one a cmov can cost a fair amount of waiting. Finally cmov were absolutely terrible on P4 (10-ish cycles), and it's likely that a lot of their lore dates back to that.
- chrisjj 2y agoThe good way is to inspect the code :) > it's also concerning that we have syntax where an if is some times a branch and some times not. It would be more concerning if we didn't. We might get a branch on one GPU and none on another.
- DrNosferatu 2y agoThis should be quantified and generalized for a full set of cases - that way the argument would stand far more clearly.
- DrNosferatu 2y agoSomething like this: https://doliveira4.github.io/gpuconditionals/ https://doliveira4.github.io/gpuconditionals/ (no warranty)
- ajross 2y ago> For the record, of course real branches do happen in GPU code Well, for some definition of "real". There are hardware features (on some architectures) that implement semantics that evaluate the same way that "branched" scalar code would. There is no branching at the instruction level, and can't be on SIMD (because the other parallel shaders being evaluated by the same instructions might not have taken the same branch!)
- account42 2y agoThere is real branching on the hardware level, but yes it needs to take the same branch for the whole workgroup and anything else needs to be "faked" in some form.
- ajross 2y agoYeah, exactly: I argue that a global backwards-only branch used to implement loops by checking all lanes for an "end" state is not actually a "real" branch. It's a semantic argument, but IMHO an important one. Way, way too many users of GPUs don't understand how the code generation actually works, leading to articles like this one. Ambiguous use of terms like "branch" are the problem.
- cjbgkagh 2y agoI think the core problem is that when writing code like this you need experience be sure that it won’t have a conditional branch. How many operations past the conditional cause a branch? Which operations can the compiler elide to bring the total below this count? I’m all for writing direct code and relying on smart compilers but it’s often hard to know if and where I’m going to get bitten. Do I always have to inspect the assembly? Do I need a performance testing suit to check for accidental regressions? I find it much easier if I can give the compiler a hint on what I expect it to do, this would be similar to a @tailcall annotation. That way I can explore the design space without worry that I’ll accidentally overstep a some hard to reason about boundary that will tank the performance.
- layer8 2y agoThis article is also relevant: https://medium.com/@jasonbooth_86226/branching-on-a-gpu-18bfc83694f2 https://medium.com/@jasonbooth_86226/branching-on-a-gpu-18bf... “If you consult the internet about writing a branch of a GPU, you might think they open the gates of hell and let demons in. They will say you should avoid them at all costs, and that you can avoid them by using the ternary operator or step() and other silly math tricks. Most of this advice is outdated at best, or just plain wrong. Let’s correct that.”
- CountHackulus 2y agoI love seeing the codegen output, makes it easy to understand the issue, but claiming that it's faster or slower without actual benchmarks is a bit disappointing.
- leeoniya 2y agothis. why waste brain cells on theory when you should simply bench both versions and validate without buying into any kind of micro-optimization advice at face value.
- grumpy_coder 2y agoI believe the conclusion is correct in 2025, but the article in a way just perpetuates the 'misinformation', making it seem like finding if your code will compile to a dynamic branch or not is easier than it is. The unfortunate truth with shaders is that they are compiled by the users machine at the point of use. So compiling it on just your machine isn't nearly good enough. NVIDIA pricing means large numbers of customers are running 10 year old hardware. Depending on target market you might even want the code to run on 10 year old integrated graphics. Does 10 year old integrated graphics across the range of drivers people actually have running prefer conditional moves over more arithmetic ops.. probably, but I would want to keep both versions around and test on real user hardware if this shader was used a lot.
- aappleby 2y agoThese sort of avoid-branches optimizations were effective once upon a time as I profiled them on the XBox 360 and some ancient Intel iGPUs, but yeah - don't do this anymore. Same story for bit extraction and other integer ops - we used to emulate them with float math because it was faster, but now every GPU has fast integer ops.
- Agentlien 2y ago> now every GPU has fast integer ops. Is that true and to what extent? Looking at the ISA for RDNA2[0] for instance - which is the architecture of both PS5 and Xbox Series S|X - all I can find is 32-bit scalar instructions for integers. [0] https://www.amd.com/content/dam/amd/en/documents/radeon-tech-docs/instruction-set-architectures/rdna2-shader-instruction-set-architecture.pdf https://www.amd.com/content/dam/amd/en/documents/radeon-tech...
- LegionMammal978 2y agoYou're likely going to have a rough time with 64-bit arithmetic in any GPU. (At least on Nvidia GPUs, the instruction set doesn't give you anything but a 32-bit add-with-carry to help.) But my understanding is that a lot of the arithmetic hardware used for 53-bit double-precision ops can also be used for 32-bit integer ops, which hasn't always been the case.
- Agentlien 2y agoI'm less concerned about it being 32-bit and more about them being exclusively scalar instructions, no vector instructions. Meaning only useful for uniforms, not thread-specific data. [Update: I remembered and double checked. While there are only scalar 32-bit integer instructions you can use 24-bit integer vector instructions. Essentially ignoring the exponent part of the floats.]
- ryao 2y agoThe programming model is that all threads in the warp / thread block run the same instruction (barring masking for branch divergence). Having SIMD instructions at the thread level is a rarity given that the way SIMD is implemented is across warps / thread blocks (groups of warps). It does exist, but only within 32-bit words and really only for limited use cases, since the proper way to do SIMD on the GPU is by having all of the threads execute the same instruction: https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#simd-video-instructions https://docs.nvidia.com/cuda/parallel-thread-execution/index... Note that I am using the Nvidia PTX documentation here. I have barely looked at the AMD RDNA documentation, so I cannot cite it without doing a bunch of reading.
- mgaunard 2y ago"of course real branches happen in GPU code" My understanding was that they don't. All executions inside a "branch" always get executed, they're simply predicated to do nothing if the condition to enter is not true.
- ack_complete 2y agoThat's only if execution is incoherent. If all threads in a warp follow the branch the same way, then all of the instructions in the not taken branch are skipped.
- yahya_6666608 2y ago[flagged]
- arbitrandomuser 2y agoWhat is the AMD and Microsoft cshader compiler , how do I generate and inpect these intermediate codes on my computer?
- blackle 2y agoFor AMD you can use the Radeon GPU Analyzer: https://gpuopen.com/rga/ https://gpuopen.com/rga/
- qwery 2y agoSome of the mistakes/confusion being pointed out in the article is being replicated here, it seems. The article is not claiming that conditional branches are free. In fact, the article is not making any point about the performance cost of branching code, as far as I can tell. The article is pointing out that conditional logic in the form presented does not get compiled into conditionally branching code. And that people should not continue to propagate the harmful advice to cover up every conditional thing in sight[0]. Finally, on actually branching code: that branching code is more complicated to execute is self-evident. There are no free branches. Avoiding branches is likely (within reason) to make any code run faster. Luckily[1], the original code was already branchless. As always, there is no universal metric to say whether optimisation is worthwhile. [0] the "in sight" is important -- there's no interest in the generated code, just in the source code not appearing to include conditional anythings. [1] No luck involved, of course ... (I assume people wrote to IQ to suggest apparently glaringly obvious (and wrong) improvements to their shader code, lol)
- cwillu 2y agoHmm, godbolt is showing branches in the vulkan output: return x>0.923880?vec2(s.x,0.0): x>0.382683?s*sqrt(0.5): vec2(0.0,s.y); turns into %24 = OpLoad %float %x %27 = OpFOrdGreaterThan %bool %24 %float_0_923879981 OpSelectionMerge %30 None OpBranchConditional %27 %29 %35 %29 = OpLabel %31 = OpAccessChain %_ptr_Function_float %s %uint_0 %32 = OpLoad %float %31 %34 = OpCompositeConstruct %v2float %32 %float_0 OpStore %28 %34 OpBranch %30 %35 = OpLabel %36 = OpLoad %float %x %38 = OpFOrdGreaterThan %bool %36 %float_0_382683009 OpSelectionMerge %41 None OpBranchConditional %38 %40 %45 %40 = OpLabel %42 = OpLoad %v2float %s %44 = OpVectorTimesScalar %v2float %42 %float_0_707106769 OpStore %39 %44 OpBranch %41 %45 = OpLabel %47 = OpAccessChain %_ptr_Function_float %s %uint_1 %48 = OpLoad %float %47 %49 = OpCompositeConstruct %v2float %float_0 %48 OpStore %39 %49 OpBranch %41 %41 = OpLabel %50 = OpLoad %v2float %39 OpStore %28 %50 OpBranch %30 %30 = OpLabel %51 = OpLoad %v2float %28 OpReturnValue %51 https://godbolt.org/z/aqob7YfWq https://godbolt.org/z/aqob7YfWq
- SideQuark 2y agoVulcan opcode shader lang is not executed. It’s a platform neutral intermediate language, so won’t have the special purpose optional instructions most GPUs do since GPUs aren’t required to. It likely compiles down on the relevant platforms as the original article did.
- nosferalatu123 2y agoA lot of the myth that "branches are slow on GPUs" is because, way back on the PlayStation 3, they were quite slow. NVIDIA's RSX GPU was on the PS3; it was documented that it was six cycles IIRC, but it always measured slower than that to me. That was for even a completely coherent branch, where all threads in the warp took the same path. Incoherent branches were slower because the IFEH instruction took six cycles, and the GPU would have to execute both sides of the branch. I believe that was the origin of the "branches are slow on GPUs" myth that continues to this day. Nowadays GPU branching is quite cheap especially coherent branches.
- nice_byte 2y agocoherent branches are "free" but the extra instructions increase register pressure. that's the main reason why dynamic branches are avoided, not that they are inherently "slow".
- dahart 2y agoIf someone says branching without qualification, I have to assume it’s incoherent. The branching mechanics might have lower overhead today, but the basic physics of the situation is that throughput on each side of the branch is reduced to the percentage of active threads. If both sides of a branch are taken, and both sides are the same instruction length, the average perf over both sides is at least cut in half. This is why the belief that branches are slow on GPUs is both persistent and true. And this is why it’s worth trying harder to reformulate the problem without branching, if possible.
- torginus 2y agoI'm not going to second guess IQ, who is one of the greatest modern authorities on shaders, but I do have some counterarguments. - Due to how SIMD works, it's quite likely both paths of the conditional statement get executed, so its a wash - Most importantly, if statements look nasty on the screen. Having an if statement means a discontinuity in visuals, which means jagged and ugly pixels on the output. Of course having a step function doesnt change this, but that means the code is already in the correct form to replace it with smoothstep, which means you can interpolate between the two variations, which does look good.
- torginus 2y agoWhy does this keep getting downvoted? This is fundamentally true, and good advice borne of experience. At least somebody would care to weight in as to why they disagree?
- grg0 2y agoI am not 100% sure nor did I downvote, but it doesn't look like a counter-argument at all. > Due to how SIMD works, it's quite likely both paths of the conditional statement get executed, so its a wash It's not just quite likely, it's what IQ is showing in the disassembly. For a ternary op like this one with trivial expressions on each side, the GPU evals both and then masks the result given the result of the condition. > a step function doesnt change this, but that means the code is already in the correct form to replace it with smoothstep, which means you can interpolate between the two variations, which does look good. A smoothstep does smooth interpolation of two values. It seems unrelated to the issue in the post. step() relates to the ternary op in the sense that both can be used to express conditionals. The post explains why you wouldn't necessarily want to use step() vs ternary op. smoothstep is related to step in some sense, but not in a way that relates to the article? i.e., going from step() to smoothstep() will entirely change the semantics of the program precisely because of the 'smooth' part.
- torginus 2y agoStep vs ternary op vs if statement are equivalent - you are correct in that. What I'm saying is not about optimization or assembly - optimization doesn't matter when the end result looks bad. What I'm saying that no matter how you express it in code, abrupt transitions of values introduce aliasing (or 'edge shimmer'), which looks unpleaseant. The way you get rid of it by smoothly blending between 2 values with smoothstep for example.
- tsylba 2y agoIt's funny because I rarely seen this (wrong approach) done anywhere else but I pick it up by myself (like a lot did I presume) and still am the first to do it everytime I see the occasion, not so for optimizations (while I admit I thought it wouldn't hurt) but for the flow and natural look of it. It feels somehow more right to me to compose effect by signals interpolations rather than clear ternary branch instructions. Now I'll have to change my ways in fear of being rejected socially for this newly approved bad practice. At least in WebGPU's WGSL we have the `select` instruction that does that ternary operation hidden as a method, so there is that.
- lerp-io 2y agoI've been doing this from day 1 becausee I just assumed you are not supposed to have loops or if else blocks in your shader code. Now I know better, thanks iq ur g.
- leguminous 2y agoUsing `mix()` isn't necessarily bad. Using boolean operations like `lessThan()` is probably better than `step()`. I just tested two ways of converting linear RGB to sRGB. On AMD, they compile to the same assembly. Method 1. float linear_to_srgb(float v) { return v < 0.0031308 ? v * 12.92 : 1.055 * pow(v, 1.0 / 2.4) - 0.055; } vec3 linear_to_srgb(vec3 rgb) { return vec3(linear_to_srgb(rgb.r), linear_to_srgb(rgb.g), linear_to_srgb(rgb.b)); } Method 2: vec3 linear_to_srgb(vec3 rgb) { bvec3 cutoff = lessThan(rgb, vec3(0.0031308)); vec3 upper = vec3(1.055) * pow(rgb, vec3(1.0 / 2.4)) - vec3(0.055); vec3 lower = rgb * vec3(12.92); return mix(upper, lower, cutoff); }