4 ms·
The overlong lookup can also be written without a memory lookup as 0x10000U >> ((0x1531U >> (i*5)) & 31); On most current x86 chips this has a latency of
by pbsd 2y ago
The overlong lookup can also be written without a memory lookup as
0x10000U >> ((0x1531U >> (i*5)) & 31);
On most current x86 chips this has a latency of 3 cycles -- LEA+SHR+SHR -- which is better than an L1 cache hit almost everywhere.
- moonchild 2y agochecking for an overlong input is off the critical path, so latency is irrelevant. (tfa doesn't appear to use a branch—it should use a branch)
- teo_zero 2y agoYou can do it with one shift: 0x10880 & (24 << (i*4))
- pbsd 2y agoBrilliant. Can be further simplified to 0x10880 & (0x2240 << i).
- hairtuq 2y agoSimilarly, the eexpect table can be done in 3 instructions with with t = 0x3c783023f + i; return ((t >> (t & 63)) & 0xf0808088.