2 ms·
At least my compiler (LLVM 10) is smart enough to know that "data" isn't modified, and thus fold the addresses of the two vector loads inside the main loop into
by zwegner 7y ago
At least my compiler (LLVM 10) is smart enough to know that "data" isn't modified, and thus fold the addresses of the two vector loads inside the main loop into single instructions:
1a0: c5 fe 6f 6c 31 ff vmovdqu ymm5,YMMWORD PTR [rcx+rsi*1-0x1]
1a6: c5 fe 6f 24 31 vmovdqu ymm4,YMMWORD PTR [rcx+rsi*1]
I'd expect most modern compilers to get this right with optimizations on, but if somebody finds one that doesn't, I'd like to know. In general, I read through the disassembly a decent amount when developing this. That's how I noticed the "req += (vmask2_t)set << n;" (instead of |=) trick, which gets compiled to one lea instruction. The disassembly got a little bit hairy when I added the code to handle trailing bytes, though...
Thanks for the kind words, though! It warms my heart :)