4 ms·
It's a lot more asm instructions than the tight loop of the naive one, as it's got to track more state (working out the middle of the slice, etc)... https://go
by pixelesque 2y ago
It's a lot more asm instructions than the tight loop of the naive one, as it's got to track more state (working out the middle of the slice, etc)...
https://godbolt.org/z/917o7oT8r https://godbolt.org/z/917o7oT8r
- kardos 2y agoAha. So it could be optimized into two tight loops. The linked wikipedia on pairwise says the numpy implementation is same speed as naive
- orlp 2y agoNumpy also uses blocked pairwise summation, with a block size of 128: https://github.com/numpy/numpy/blob/a6e9dc7152098182b45ecd6effa15223890d663e/numpy/_core/src/umath/loops_utils.h.src#L55 https://github.com/numpy/numpy/blob/a6e9dc7152098182b45ecd6e... .