2 ms·
The hand-coded AVX2 procedure is far from optimal form, they waste time on horizontal addition in each iteration. The First Rule of SIMD-ization says: keep all
by wmu 7y ago
The hand-coded AVX2 procedure is far from optimal form, they waste time on horizontal addition in each iteration.
The First Rule of SIMD-ization says: keep all the intermediate results in vector(s), do horizontal reduction at the end.
Conversion from comparison result into vector of integers can be done a bit simpler: just one bit-and is needed and then cast to __m256i (casting doesn't emit any code as SIMD registers are untyped).