5 ms·
L2 speeds are ~180GB/s on Zen 4. That's also a part of my confusion. This should be line speed. I do not have the same experience as you with unrolling AVX512
by kolbe 2y ago
L2 speeds are ~180GB/s on Zen 4. That's also a part of my confusion. This should be line speed.
I do not have the same experience as you with unrolling AVX512 loops on Zen 4. I recall even with double pumping, you can do 2.5 per cycle. As you noted with the stores, it takes two cycles, so you can put 4 into a 3.25 cycle pipeline, instead of 8 cycles. With 5 dependent ops covering 4.5 cycles, this should be a significant win.
I'm not defending the 31 length loop to clean up the mod32 leftover section. That is bad. But it doesn't answer why 256B is 4x slower than line speed and not significantly faster than the unrolled intrinsic version.
In my experience, I never used the masked load store because some platforms did an actual read over the masked-away parts, and could segfault. I recall hearing from a reliable source that Zen 4 doesn't do that, but didn't see official documentation for it. Clang may actually be avoiding the masked cleanup for that reason.
To top it off, I also always found it faster to just stagger the index back instead of a masked load/store whenever it's longer than 64 on calculations like this. That is, if it's size=80, do 0-63, and then 15-79 (which is an optimization Clang doesn't do either for some reason).
Finally, what really really confuses me is that whenever I write a benchmark like this:
for(size_t i = 0; i < sizeof(src); i++) {
src[i] = (uint8_t)(i & 63) + 32;
}
-O3 will do something absurd like just return the answer, since the inputs are constexpr. I can understand why the intrinsic version might confuse the compiler, but the clearly written one should totally have broken the benchmark and overwritten it with a constexpr answer.
- dzaima 2y agoAlso of concern is that the input, and thus the loads & stores here, are intentionally bumped to cycle through all possible alignments, thus ending up unaligned most of the time, in which case it should be 3 cycles per store (I think?). I don't understand your point about pipelining - OoO should mean that, as long as there's enough decode bandwidth and per-iteration scalar overhead doesn't overwhelm scalar execution resources, all SIMD ops can run at full force up to the most contended resource (store here), no? That said, yeah, ~44GB/s is actually still pretty slow here, even for L3. Masked load/store faulting was problematic on AVX2 (in addition to being pretty slow on AMD (which actually continues into Zen 4, despite it having fast AVX-512 versions)); AVX-512's should always be fine, and compilers already output them: https://godbolt.org/z/98sY57TE1 https://godbolt.org/z/98sY57TE1 Intrinsics shouldn't "confuse" clang - clang lowers them, where possible, to the same LLVM instructions that the autovectorizer would generate. Both clang and gcc can even convert an intrinsics-based memcpy/memset impl to a libc call (as annoying may that be)! If you want a compiler to not optimize out computation, you can add something like `__asm__ volatile(""::"r"(src):"memory");` after the loop to make it operate as if the contents of src were modified/read.
- fanf2 2y agoClang rewrites my tolower64() into something rather more clever, and it adds a 2x unrolled version. (gcc keeps the object code much closer to what I wrote.) But from my point of view the important thing is that Clang retains the masked load and store for short string fragments. Its autovectorizer is not able to generate masked loads and stores. I used Clang 11 for copybytes64() because it is unable to recognize the memcpy() idiom, whereas Clang 16 does turn it into memcpy() - which is slower!
- kolbe 2y ago> I don't understand your point about pipelining - OoO should mean that, as long as there's enough decode bandwidth and per-iteration scalar overhead doesn't overwhelm scalar execution resources, all SIMD ops can run at full force up to the most contended resource (store here), no? You are reaching the limits of my understanding, but my level of knowledge is that store may have reciprocal throughput of 2, but it only occupies two ops (from double pumping a single one) over those two cycles, while the CPU pipeline can handle doing 10. For store in particular, nothing is dependent on it completing, so it can be "thrown into the wind" so to speak. But here's my approximation of the pipeline of a single thread, where dashes separate ops LOADU.0 - LOADU.1 - _ - _ - _ - _ - ADD.0 - ADD.1 - _ - _ - CMP.0 - CMP.1 - _ - _ - _ - _ - ADD.0 - ADD.1 - _ - _ - STORE.0 - STORE.1 - [start again, because nothing is dependent on STORE completing] So, that's 10 ops and 12 empty spots that can be filled by simultaneously doing 1.2 more loops simultaneously. I do want to know why clang isn't using the masked load/store. If it's willing to do it on a dot-product, it should do it here as well. It makes me want to figure out what is blocking it (usually some guarantee that 99.9% of developers don't know they're making).
- dzaima 2y agoThe store will still take up throughput even if nothing depends on it right now - there is limited hardware available for copying data from the register file to the cache, and its limit is two 32-byte stores per cycle, which you'll have to pay one way or another at some point. With out-of-order execution, the layout of instructions in the source just doesn't matter at all - the CPU will hold multiple iterations of the loop in the reorder buffer, and assign execution units from multiple iterations. e.g. see: https://uica.uops.info/?code=vmovdqu64%20zmm3%2C%20zmmword%20ptr%20%5Brsi%5D%0D%0Avpaddb%20%20%20%20zmm4%2C%20zmm3%2C%20zmm0%0D%0Avpcmpltub%20k1%2C%20zmm4%2C%20zmm1%0D%0Avpaddb%20%20%20%20zmm3%20%7Bk1%7D%2C%20zmm3%2C%20zmm2%0D%0Avmovdqu64%20zmmword%20ptr%20%5Brdi%5D%2C%20zmm3&syntax=asIntel&uArchs=TGL&tools=uiCA&alignment=0&uiCAHtmlOptions=traceTable https://uica.uops.info/?code=vmovdqu64%20zmm3%2C%20zmmword%2... (click run, then Open Trace); That's Tiger Lake, not Zen 4, but still displays how instructions from multiple iterations execute in parallel. Zen 4's double-pumping doesn't change the big picture, only essentially meaning that each zmm instr is split into two ymm ones (they might not even need to be on the same port, i.e. double-pumping is really the wrong term, but whatever).
- fanf2 2y agoThe benchmark measures the time to copy about 1 MiByte, in chunks of various lengths from 1 byte to 1 kilobyte. I wanted to take into account differences in alignment in the source and destination strings, so there are a few bytes between each source and destination string, which are not counted as part of the megabyte. On my Zen 4 CPU the L2 cache is 1 MiB per core, so because the working set is over 2 MB I think the relevant speed limit is the L3 cache. To be sure I was measuring what I thought I was, I compiled each function separately to avoid interference from inlining, code motion, loop fusion, etc. Of course in real code it's more likely that you would want to encourage inlining, not prevent it!
- IWeldMelons 2y agoNever heard about reading the masked area. Source?
- dzaima 2y agoAMD64 manual volume 4 (https://www.amd.com/content/dam/amd/en/documents/processor-tech-docs/programmer-references/26568.pdf https://www.amd.com/content/dam/amd/en/documents/processor-t...): > Exception and trap behavior for elements not selected for loading or storing from/to memory is implementation dependent. For instance, a given implementation may signal a data breakpoint or a page fault for doublewords that are zero-masked and not actually written.
- IWeldMelons 2y agoI wonder if it ever happens IRL, as this would severely restrict the usefulness. EDIT: does not seem to be applicable to AVX-512, only to AVX 1. According to this PDF by intel, AVX-512 suppresses faults for masked access. https://gcc.gnu.org/wiki/cauldron2014?action=AttachFile&do=get&target=Cauldron14_AVX-512_Vector_ISA_Kirill_Yukhin_20140711.pdf https://gcc.gnu.org/wiki/cauldron2014?action=AttachFile&do=g...
- IWeldMelons 2y agoAccording to Intel document https://gcc.gnu.org/wiki/cauldron2014?action=AttachFile&do=get&target=Cauldron14_AVX-512_Vector_ISA_Kirill_Yukhin_20140711.pdf https://gcc.gnu.org/wiki/cauldron2014?action=AttachFile&do=g..., AVX-512 should suppress faults during masked access or otherwise it is not AVX-512.