5 ms·
The main point is that the conditional didn't actually introduce a branch. Showing the other generated version would only show that it's longer. It is not expe
by azeemba 2y ago
The main point is that the conditional didn't actually introduce a branch.
Showing the other generated version would only show that it's longer. It is not expected to have a branch either. So I don't think it would have added much value
- idunnoman1222 2y agoUnless you’re writing an essay on why you’re right…
- chrisjj 2y ago> Unless you’re writing an essay on why you’re right… He's writing an essay on why they are wrong. "But here's the problem - when seeing code like this, somebody somewhere will invariably propose the following "optimization", which replaces what they believe (erroneously) are "conditional branches" by arithmetical operations." Hence his branchless codegen samples are sufficient. Further, regarding.the side-issue "The second wrong thing with the supposedly optimizer [sic] version is that it actually runs much slower", no amount of codegen is going to show lower /speed/.
- ncruces 2y agoThe other either optimizes the same, or has an additional multiplication, and it's definitely less readable.
- TheRealPomax 2y agoCorrect: it would show proof instead of leaving it up to the reader to believe them.
- comex 2y agoBut it's possible that the compiler is smart enough to optimize the step() version down to the same code as the conditional version. If true, that still wouldn't justify using step(), but it would mean that the step() version isn't "wasting two multiplications and one or two additions" as the post says. (I don't know enough about GPU compilers to say whether they implement such an optimization, but if step() abuse is as popular as the post says, then they probably should.)
- MindSpunk 2y agoOkay but how does this help the reader? If the worse code happens to optimize to the same thing it's still awful and you get no benefits. It's likely not to optimize down unless you have fast-math enabled because the extra float ops have to be preserved to be IEEE754 compliant
- burnished 2y ago..how is it awful if it has the same result?
- account42 2y agoFragment and vertex shaders generally don't target strict IEEE754 compliance by default. Transforming a * (b ? 1.0 : 0.0) into b ? a : 0.0 is absolutely something you can expect a shader compiler to do - that only requires assuming a is not NaN.
- Lockal 2y agoYou missed the second part where article says that "it actually runs much slower than the original version", "wasting two multiplications and one or two additions", based on idea that compiler is unable to do a very basic optimization, implying that compiler compiler will actually multiply by one. No benchmarks, no checking assembly, just straightforward misinformation.