6 ms·
Well, the conundrum sits with the fact that this is the disassembly of the main loop: vmovdqu64 zmm3, zmmword ptr [rdi + rcx] vmovdqu64
by kolbe 2y ago
Well, the conundrum sits with the fact that this is the disassembly of the main loop:
vmovdqu64 zmm3, zmmword ptr [rdi + rcx]
vmovdqu64 zmm4, zmmword ptr [rdi + rcx + 64]
vmovdqu64 zmm5, zmmword ptr [rdi + rcx + 128]
vmovdqu64 zmm6, zmmword ptr [rdi + rcx + 192]
vpaddb zmm7, zmm3, zmm0
vpaddb zmm8, zmm4, zmm0
vpaddb zmm9, zmm5, zmm0
vpaddb zmm10, zmm6, zmm0
vpcmpltub k1, zmm7, zmm1
vpcmpltub k2, zmm8, zmm1
vpcmpltub k3, zmm9, zmm1
vpcmpltub k4, zmm10, zmm1
vpaddb zmm3 {k1}, zmm3, zmm2
vpaddb zmm4 {k2}, zmm4, zmm2
vpaddb zmm5 {k3}, zmm5, zmm2
vpaddb zmm6 {k4}, zmm6, zmm2
vmovdqu64 zmmword ptr [rdi + rcx], zmm3
vmovdqu64 zmmword ptr [rdi + rcx + 64], zmm4
vmovdqu64 zmmword ptr [rdi + rcx + 128], zmm5
vmovdqu64 zmmword ptr [rdi + rcx + 192], zmm6
...which is, upon first glance, is similar to yet better than the intrinsics version you wrote. Additionally it has cleaner tail handling.
- fanf2 2y agoThe tail handling is the main point of the post. The tolower1 line on the chart is very spiky because the autovectorizer doesn’t use masked loads and stores for tails, and instead does something slower that tanks performance. The tolower64 line is smoother and rises faster because masked loads and stores make it easier to handle strings shorter than the vector size.
- kolbe 2y agoThere is something to be said for using masking load/stores instead of downgrading the SIMD register types. But to my eye, something is clearly wrong with your code in such a way that the "naive" autovectorized version should be blowing yours out of the water at 256B. Your code leaves a sparse instruction pipeline, and thus needs to run through the loop 4 times for every one of the autovectorized version. So, something is wrong here, as far as I know. I was just trying to make sure you'd covered the basics, like ensuring you were targeting the right architecture with your build and whatnot. This is yours. Basically the same instructions, but taking up way more space: vmovdqu64 zmm3, zmmword ptr [rsi] vpaddb zmm4, zmm3, zmm0 vpcmpltub k1, zmm4, zmm1 vpaddb zmm3 {k1}, zmm3, zmm2 vmovdqu64 zmmword ptr [rdi], zmm3
- dzaima 2y agoUnrolling on AVX-512, especially on Zen 4 with its double-pumped almost-everything, isn't particularly significant; the 512-bit store alone has a reciprocal throughput of 2 cycles/zmm, which gives a pretty comfortable best possible target of 2 cycles/iteration. With Zen 4 being 6-wide out of the op cache, that's fine for up to 12 instructions in the loop (and with separate ports for scalar & SIMD, the per-iteration scalar op overhead is irrelevant). Interestingly enough, clang 16 doesn't unroll the intrinsics version, but trunk does, which'd make the entire point moot. The benchmark in question, as per the article, tests over 1MB in blocks, so it'll be at L2 or L3 speeds anyway. Clang downgrades to ymm, which'll handle a 32-byte tail, but after that it does a plain scalar loop, up to a massive 31 iterations. Whereas masking appears to be approximately free on Zen 4, so I'd be surprised if the scalar tail would be better at even, like, 4 bytes (perhaps there could be some minor throughput gained by splitting up into optional unmasked ymm and then a masked ymm, but even that's probably questionable just from the extra logic/decode overhead). Also worth considering that in real-world usage the masking version could be significantly better just from being branchless for up to 64-byte inputs.
- fanf2 2y agoYeah the spikes in the “tolower1” line illustrate the 32 byte tail pretty nicely. I should maybe draw a version of the chart covering just small strings, but it’s an SVG so you can zoom right in. The “tolower1” line shows relatively poor performance compared to “tolower64”, tho it is hard to see for strings less than 8 bytes.
- kolbe 2y agoL2 speeds are ~180GB/s on Zen 4. That's also a part of my confusion. This should be line speed. I do not have the same experience as you with unrolling AVX512 loops on Zen 4. I recall even with double pumping, you can do 2.5 per cycle. As you noted with the stores, it takes two cycles, so you can put 4 into a 3.25 cycle pipeline, instead of 8 cycles. With 5 dependent ops covering 4.5 cycles, this should be a significant win. I'm not defending the 31 length loop to clean up the mod32 leftover section. That is bad. But it doesn't answer why 256B is 4x slower than line speed and not significantly faster than the unrolled intrinsic version. In my experience, I never used the masked load store because some platforms did an actual read over the masked-away parts, and could segfault. I recall hearing from a reliable source that Zen 4 doesn't do that, but didn't see official documentation for it. Clang may actually be avoiding the masked cleanup for that reason. To top it off, I also always found it faster to just stagger the index back instead of a masked load/store whenever it's longer than 64 on calculations like this. That is, if it's size=80, do 0-63, and then 15-79 (which is an optimization Clang doesn't do either for some reason). Finally, what really really confuses me is that whenever I write a benchmark like this: for(size_t i = 0; i < sizeof(src); i++) { src[i] = (uint8_t)(i & 63) + 32; } -O3 will do something absurd like just return the answer, since the inputs are constexpr. I can understand why the intrinsic version might confuse the compiler, but the clearly written one should totally have broken the benchmark and overwritten it with a constexpr answer.
- harry8 2y agoThe autovectorizer can, at times, when you check it produce reasonable simd code with a given compiler based on simple, clear C code. Is gcc as good in this case? Are previous versions of clang (that may be employed by users) going to work out as well. How do you check in your build that the compiler did the right thing? If you touch the C code in any way will it silently do something you don't want? Autovectorization is great and improving.