5 ms·
This article seems to be written very unclearly, but by "reduction", it seems be means a vectorized function that takes a double[4] and sums the 4 values, and f
by BorisTheBrave 6y ago
This article seems to be written very unclearly, but by "reduction", it seems be means a vectorized function that takes a double[4] and sums the 4 values, and fills another double[4] with 4 copies of the result.
I don't really see any explanation of why you'd want to do this? Not only is there a instruction for doing exactly this in most vector instruction sets, but you typically don't need to do it inside your hot loop anyway.
- tom_mellior 6y ago> by "reduction", it seems be means a vectorized function that takes a double[4] and sums the 4 values "Reduce" or "fold" is a standard term for applying a binary operation (like +) to all elements of a vector/list/more general data structure: https://en.wikipedia.org/wiki/Fold_(higher-order_function) https://en.wikipedia.org/wiki/Fold_(higher-order_function) > Not only is there a instruction for doing exactly this in most vector instruction sets Is there a single one in AVX for adding all the components of a vector register? Is it fast? How does it associate the computation? If there is such an instruction, GCC and Clang seem reluctant to emit it: https://gcc.godbolt.org/z/vkaeYi https://gcc.godbolt.org/z/vkaeYi though I might be missing some magic flags.
- gpderetta 6y agoUnless I completely misunderstood, isn't that an horizontal add, ie. HADDPD[1]? edit: apparently GCC doesn't specifically optimize reductions [2] and does a generic vectorization instead. [1] https://www.felixcloutier.com/x86/haddpd https://www.felixcloutier.com/x86/haddpd [2] https://gcc.gnu.org/bugzilla/show_bug.cgi?id=54400 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=54400
- tom_mellior 6y agoYes this is a horizontal add, but HADDPD only does a very specific part of a full horizontal add. You can use it to compute x[0] + x[1] in part of a register and x[2] + x[3] in another part, but afterwards you would still need to shuffle and add.
- zbjornson 6y agoIt's also slow (5-6 cycle latency and 2 cycle throughput on Intel processors, compared to 4/0.5 for multiplication).
- gnufx 6y agoGCC 8 -Ofast, targeting Haswell, produces vhaddpd %xmm0, %xmm0, %xmm0
- tom_mellior 6y agoDo you have an input program and an exact command line to experiment with? Using my function from above with -Ofast -ffast-math -march=haswell does not produce this on the GCC 8.1 on Compiler Explorer.
- gnufx 6y agoJust per #54400 above. $ gcc --version | head -1 gcc (Debian 8.3.0-6) 8.3.0 $ cat x.c #include <x86intrin.h> double f(__m128d v){return v[1]+v[0];} $ gcc -Ofast -march=haswell -S x.c $ grep -C1 hadd x.s .cfi_startproc vhaddpd %xmm0, %xmm0, %xmm0 ret GCC 4.8.5 with adjusted options does the same, and 7.5, so perhaps there was some regression in early 8. (-ffast-math is redundant with -Ofast.)
- tom_mellior 6y agoThanks. There probably wasn't a regression, it's just that this is not the code I was asking about. Your code works for the special case of horizontally adding two elements, but not for the more general case (4 elements) that this whole thread was about.
- gpderetta 6y agooh, interesting, I explicitly tried -ffast-math and a few optimization targets, with not haswell explicitly.