4 ms·
Bit-fiddling at the C-level isn't necessarily faster on all platforms. E.g., the "masked combine bits" case where he says that r = a ^ ((a ^ b) & mask) is one
by reichstein 15y ago
Bit-fiddling at the C-level isn't necessarily faster on all platforms.
E.g., the "masked combine bits" case where he says that
r = a ^ ((a ^ b) & mask)
is one operation shorter than
r = (a ^ ~mask) | (b ^ mask).
On ARM, the (a ^ ~mask) can be performed by one instruction, so the instruction count is actually the same, and the latter way of doing it parallelizes better (only a dependency depth of two instead of three).
I.e., the "optimized" version is actually likely to be slower on some CPUs.