4 ms·
Does anyone have any tips for similar wizardry-level SIMD optimization on ARM?
by lumb63 2y ago
Does anyone have any tips for similar wizardry-level SIMD optimization on ARM?
- anonymoushn 2y agoIf you learn AVX2 programming via highload, my impression is that NEON is quite similar. The main difference is the lack of movemask. You can read these[0] articles[1] about what to do instead. For SVE, prior to very recent versions of SVE, there was no tblq (pshufb equivalent) so I didn't have much hope for using it for general-purpose programming, though of course it would be fine for stuff like TFA. [0]: https://community.arm.com/arm-community-blogs/b/infrastructure-solutions-blog/posts/porting-x86-vector-bitmask-optimizations-to-arm-neon https://community.arm.com/arm-community-blogs/b/infrastructu... [1]: https://www.corsix.org/content/whirlwind-tour-aarch64-vector-instructions https://www.corsix.org/content/whirlwind-tour-aarch64-vector...
- TinkersW 2y agoIt isn't as wide on ARM so the gains will be smaller(most ARM is only 128 bit wide neon)