4 ms·
No, I think they're both using memcpy. perf says they're spending all the time in libc at least. The total time is a little slower than operations like rotate t
by mlochbaum 4y ago
No, I think they're both using memcpy. perf says they're spending all the time in libc at least. The total time is a little slower than operations like rotate that I know are using memcpy (0.18s), so the chunk size does seem to introduce some overhead.
- moonchild 4y agoGlibc strings functions are ok, but not great (but one-size-fits-all is hard, and for all I know they could be fine here). While 512 bytes seems like a lot, the core loop is likely 4x unrolled avx; that is 128 bytes per iteration, so only 4 iterations. And it probably tries to align the dst, so annoying dispatch overhead (which you could avoid by working in batches of k at a time, eg k=4 for double floats and avx in the general case).
- mlochbaum 4y agoThe big difference is the access pattern: see the benchmarks below. Index does the small memcpys, but it speeds up if the indices are in order (my earlier benchmark used ⌽, not ⊖, because I don't remember APL any more). So prefetching might help. I guess it's possible that a larger-scale blocking would too? But 5GB/s for ⊖ isn't great either. In an application that uses these huge arrays and needs the best performance (most don't!), it needs to be split up, ideally into sections that fit in L1, so that multiple array operations can be applied to those chunks and stay CPU-bound (at least for transpose, maybe not for arithmetic). That's why I wouldn't be too interested in optimizing this case. ⎕IO←0 ⋄ ar←,[0 1 2]a←?27 1000 40 77⍴0 ci←(⊢⍴∘⍳×/)3↑⍴a ⍝ cell indices i←⍉ci ⋄ cmpx '2 1 0 3⍉a' 'ar[i;]' 2 1 0 3⍉a → 2.8E¯1 | 0% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕ ar[i;] → 2.8E¯1 | 0% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕ i←5⊖ci ⋄ cmpx '5⊖a' 'ar[i;]' 5⊖a → 1.2E¯1 | 0% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕ ar[i;] → 1.7E¯1 | +46% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕