3 ms·
The big difference is the access pattern: see the benchmarks below. Index does the small memcpys, but it speeds up if the indices are in order (my earlier bench
by mlochbaum 4y ago
The big difference is the access pattern: see the benchmarks below. Index does the small memcpys, but it speeds up if the indices are in order (my earlier benchmark used ⌽, not ⊖, because I don't remember APL any more). So prefetching might help. I guess it's possible that a larger-scale blocking would too?
But 5GB/s for ⊖ isn't great either. In an application that uses these huge arrays and needs the best performance (most don't!), it needs to be split up, ideally into sections that fit in L1, so that multiple array operations can be applied to those chunks and stay CPU-bound (at least for transpose, maybe not for arithmetic). That's why I wouldn't be too interested in optimizing this case.
⎕IO←0 ⋄ ar←,[0 1 2]a←?27 1000 40 77⍴0
ci←(⊢⍴∘⍳×/)3↑⍴a ⍝ cell indices
i←⍉ci ⋄ cmpx '2 1 0 3⍉a' 'ar[i;]'
2 1 0 3⍉a → 2.8E¯1 | 0% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕
ar[i;] → 2.8E¯1 | 0% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕
i←5⊖ci ⋄ cmpx '5⊖a' 'ar[i;]'
5⊖a → 1.2E¯1 | 0% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕
ar[i;] → 1.7E¯1 | +46% ⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕⎕