5 ms·
Did you try something like this? The autovectorizer looks pretty clean to me. https://godbolt.org/z/1c5joKK5n https://godbolt.org/z/1c5joKK5n
by kolbe 2y ago
Did you try something like this? The autovectorizer looks pretty clean to me.
https://godbolt.org/z/1c5joKK5n https://godbolt.org/z/1c5joKK5n
- fanf2 2y agoThat’s basically the same as `tolower1` - see the bullet points below the graph
- kolbe 2y agoWell, the conundrum sits with the fact that this is the disassembly of the main loop: vmovdqu64 zmm3, zmmword ptr [rdi + rcx] vmovdqu64 zmm4, zmmword ptr [rdi + rcx + 64] vmovdqu64 zmm5, zmmword ptr [rdi + rcx + 128] vmovdqu64 zmm6, zmmword ptr [rdi + rcx + 192] vpaddb zmm7, zmm3, zmm0 vpaddb zmm8, zmm4, zmm0 vpaddb zmm9, zmm5, zmm0 vpaddb zmm10, zmm6, zmm0 vpcmpltub k1, zmm7, zmm1 vpcmpltub k2, zmm8, zmm1 vpcmpltub k3, zmm9, zmm1 vpcmpltub k4, zmm10, zmm1 vpaddb zmm3 {k1}, zmm3, zmm2 vpaddb zmm4 {k2}, zmm4, zmm2 vpaddb zmm5 {k3}, zmm5, zmm2 vpaddb zmm6 {k4}, zmm6, zmm2 vmovdqu64 zmmword ptr [rdi + rcx], zmm3 vmovdqu64 zmmword ptr [rdi + rcx + 64], zmm4 vmovdqu64 zmmword ptr [rdi + rcx + 128], zmm5 vmovdqu64 zmmword ptr [rdi + rcx + 192], zmm6 ...which is, upon first glance, is similar to yet better than the intrinsics version you wrote. Additionally it has cleaner tail handling.
- fanf2 2y agoThe tail handling is the main point of the post. The tolower1 line on the chart is very spiky because the autovectorizer doesn’t use masked loads and stores for tails, and instead does something slower that tanks performance. The tolower64 line is smoother and rises faster because masked loads and stores make it easier to handle strings shorter than the vector size.
- kolbe 2y agoThere is something to be said for using masking load/stores instead of downgrading the SIMD register types. But to my eye, something is clearly wrong with your code in such a way that the "naive" autovectorized version should be blowing yours out of the water at 256B. Your code leaves a sparse instruction pipeline, and thus needs to run through the loop 4 times for every one of the autovectorized version. So, something is wrong here, as far as I know. I was just trying to make sure you'd covered the basics, like ensuring you were targeting the right architecture with your build and whatnot. This is yours. Basically the same instructions, but taking up way more space: vmovdqu64 zmm3, zmmword ptr [rsi] vpaddb zmm4, zmm3, zmm0 vpcmpltub k1, zmm4, zmm1 vpaddb zmm3 {k1}, zmm3, zmm2 vmovdqu64 zmmword ptr [rdi], zmm3
- dzaima 2y agoUnrolling on AVX-512, especially on Zen 4 with its double-pumped almost-everything, isn't particularly significant; the 512-bit store alone has a reciprocal throughput of 2 cycles/zmm, which gives a pretty comfortable best possible target of 2 cycles/iteration. With Zen 4 being 6-wide out of the op cache, that's fine for up to 12 instructions in the loop (and with separate ports for scalar & SIMD, the per-iteration scalar op overhead is irrelevant). Interestingly enough, clang 16 doesn't unroll the intrinsics version, but trunk does, which'd make the entire point moot. The benchmark in question, as per the article, tests over 1MB in blocks, so it'll be at L2 or L3 speeds anyway. Clang downgrades to ymm, which'll handle a 32-byte tail, but after that it does a plain scalar loop, up to a massive 31 iterations. Whereas masking appears to be approximately free on Zen 4, so I'd be surprised if the scalar tail would be better at even, like, 4 bytes (perhaps there could be some minor throughput gained by splitting up into optional unmasked ymm and then a masked ymm, but even that's probably questionable just from the extra logic/decode overhead). Also worth considering that in real-world usage the masking version could be significantly better just from being branchless for up to 64-byte inputs.
- fanf2 2y agoYeah the spikes in the “tolower1” line illustrate the 32 byte tail pretty nicely. I should maybe draw a version of the chart covering just small strings, but it’s an SVG so you can zoom right in. The “tolower1” line shows relatively poor performance compared to “tolower64”, tho it is hard to see for strings less than 8 bytes.
- harry8 2y agoThe autovectorizer can, at times, when you check it produce reasonable simd code with a given compiler based on simple, clear C code. Is gcc as good in this case? Are previous versions of clang (that may be employed by users) going to work out as well. How do you check in your build that the compiler did the right thing? If you touch the C code in any way will it silently do something you don't want? Autovectorization is great and improving.