4 ms·
Wow, not bad! I'm looking forward to the writeup. I'm still thinking about this one. It seems that ~.5 cycles/byte is slower than it should be. There's 6 pext-
by zwegner 8y ago
Wow, not bad! I'm looking forward to the writeup.
I'm still thinking about this one. It seems that ~.5 cycles/byte is slower than it should be. There's 6 pext->kmovq->vpaddb chains, where the pexts and kmovqs are all independent and can be pipelined. One chain should be 3+2+1 cycles, so all 6 should be 11 cycles. Agner Fog's site doesn't have information for vpermb yet, but a Stack Overflow comment says it's 1 cycle (which is rather surprising, that's 64 64->1 8 bit muxes--lots of silicon!). The latency on most other instructions can be hidden.
I updated my code with a couple minor optimizations, that at least lead to some nicer assembly output:
Getting rid of the if (mask) branch--it should almost never be false in real code (and never is in the microbenchmarks).
Inverting the mask during the comparison:
uint64_t mask = _mm512_cmpneq_epi8_mask(input, spaces)
& _mm512_cmpneq_epi8_mask(input, NL)
& _mm512_cmpneq_epi8_mask(input, CR);
...then the mask computation gets a lot cleaner, since there's now just three chained vpcmps (though this might lose a couple cycles or so because of the dependent 3 cycle latencies. Not sure how to trick gcc into using parallel vcmps with the ands on GPRs...). But we also save the mask = ~mask and the subtraction from 64 of the popcount.
All in all, speed-of-light for this code should be 3*3+2+11+1=23 cycles per iteration (if I'm counting everything properly), and probably even faster from pipelining between iterations. I hope it should be a bit faster than the 32ish that the old version was using.
- BeeOnRope 8y agovpermb is 3L1T, ie three cycles latency, one per cycle throughput. Full AIDA64 timing dump for cannonlake are on instlatx64 side.
- wmu 8y agoPerhaps nobody would see this, as the thread is already dead, but... :) I managed to include description of your algorithm and updated results, including procedures by aqrit: http://0x80.pl/notesen/2019-01-05-avx512vbmi-remove-spaces.html http://0x80.pl/notesen/2019-01-05-avx512vbmi-remove-spaces.h... BTW: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=88798 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=88798