4 ms·
Intel's intrinsics are just a total mess. Another example is _mm_loadl_epi64(), which loads 64-bits from memory and maps to MOVQ x,m -- but it takes an __m128i
by ack_complete 3y ago
Intel's intrinsics are just a total mess. Another example is _mm_loadl_epi64(), which loads 64-bits from memory and maps to MOVQ x,m -- but it takes an __m128i pointer. The result is that you frequently have to reinterpret cast pointers to use the intrinsics, in ways that would be blatantly broken in any other situation. ARM does a much better job of this in properly using void pointers or providing typed loads and store intrinsics.
It's also fun how intrinsics are mixed between ISA extension levels with confusingly similar names that make it really easy to use them in the wrong code path. Arithmetic shift right immediate for int16 (_mm_srai_epi16) and int32 (_mm_srai_epi32) are SSE2. But _mm_srai_epi64 for int64 is AVX512.
One thing that Intel does do better is that they have a standardized, OS-independent way of testing for ISA extensions through CPUID. ARM doesn't and you are at the mercy of the OS to provide you APIs to test whether, for instance, Crypto, CRC32, CAS, and UDOT instructions are available.