4 ms·
yes but >> 8 is so much faster
by lacedeconstruct 4mo ago
yes but >> 8 is so much faster
- dist-epoch 4mo agoOnly in micro-benchmarks. For real usage, today's CPUs are limited by memory bandwidth.
- szundi 4mo ago[dead]
- lacedeconstruct 4mo agoWhat are you talking about in a hot loop in my software renderer this is like 10x faster // color4_t result = { // .r = (src.r * src.a + dst.r * inv_alpha) * INV_255, // .g = (src.g * src.a + dst.g * inv_alpha) * INV_255, // .b = (src.b * src.a + dst.b * inv_alpha) * INV_255, // .a = src.a + (dst.a * inv_alpha) * INV_255 // }; // 1/256 but much faster color4_t result = { .r = (src.r * src.a + dst.r * inv_alpha) >> 8, .g = (src.g * src.a + dst.g * inv_alpha) >> 8, .b = (src.b * src.a + dst.b * inv_alpha) >> 8, .a = src.a + ((dst.a * inv_alpha) >> 8) };
- dist-epoch 4mo agoBecause you are working in the cache. Also, you should use SIMD.
- lacedeconstruct 4mo ago> Also, you should use SIMD. ironically no clang is better at auto vectorizing
- spider-mario 4mo agoBetter than what? And do you use `-mavx2` or do you let it target baseline x86_64 and miss out on 8-float vectors? How do you make sure its autovectorisation is successful?
- Tuna-Fish 4mo agoIf the latter is 10x faster, the issue is some kind of weird compilation failure for the above version. For one, it only cuts a third of the multiplies.
- virtualritz 4mo agoAnd both are wrong since the values would have to be in a linear color space for for the compositing math to make sense. But in some non-linear space to be useful when mapped to 0..255 (e.g non-linear sRGB). Which happens right after the Porter-Duff Over operator above -- a smoking gun. Which one is it gonna be? I.e. the display transform is omitted from this and the math involved with the latter makes your whole argument moot. It can't be expressed well enough with bitshifts to keep your purported 10x speedup anyway (and which I strongly doubt btw). And lastly: in a software renderer that stuff is usually <0.01% of the compute in the absolut worst case. P.S.: I'm speaking from 30 years of experience with software rendering in the context of VFX.
- imtringued 4mo agoHow is this supposed to be 10x faster if all you did was drop one out of three multiplications?
- StilesCrisis 4mo agoIt's just multiplication. Floating multiply is extraordinarily fast.
- lacedeconstruct 4mo agoThe difference between 20 cycles and 1 clock cycle in a hot loop is very noticeable
- Sesse__ 4mo agoUseful, then, that you can start several vectorized floating-point muls each cycle. (E.g., most modern x86 are 3/0.5 cycles for vmulps. No 20 cycles in sight.)
- Tuna-Fish 4mo agoFP Division by constant is optimized by a compiler into a multiply. Graphics processing typically happens on the GPU these days, and on all recent GPUs FPMUL belongs to the class of lowest-latency operations. That is, there are no other instructions that complete faster.
- mgaunard 4mo agoThat's only valid to do if the reciprocal is representable exactly.
- hansvm 4mo agoThat's not totally true. It's sufficient to be exactly representable, but you only need the reciprocal rounding error to be small enough to guarantee the multiplication rounding step fixes it across the entire range of numerators. For IEEE754 f16 values, there are 28 such extra values, the positive and negative sides of 1705/x where x is a power of 2 at least as great as 2048.
- mgaunard 4mo ago
- xigoi 4mo agoYou don’t divide a float by 256 by shifting it right eight bits; that would yield complete garbage. You subtract 8 from the exponent, then check if you got an underflow.
- dheera 4mo agoSame point; divide by power of 2 is a fast subtraction operation in float world, while divide by 255 shits all over the whole float
- yongjik 4mo agoIf your input is an arbitrary float, you need to check for denormals (and maybe NaNs). You can do bitmasking trick to avoid conditional jumps but I'm skeptical you can do it faster than SIMD multiply instruction.