3 ms·
The current PR for ARM SIMD[1] uses a different instruction mix to achieve the same goals as movemask. I tested the PR and it has a significant speedup over the
by celeritascelery 3y ago
The current PR for ARM SIMD[1] uses a different instruction mix to achieve the same goals as movemask. I tested the PR and it has a significant speedup over the non-vectorized version.
[1]https://github.com/BurntSushi/memchr/pull/114 https://github.com/BurntSushi/memchr/pull/114
- burntsushi 3y agoYup, thanks for the reminder. That's my starting point once my M2 arrives.