3 ms·
You run into these problems on x86, also, with SIMD intrinsics like _mm_cvtepi8_epi32() (which converts four 8-bit integers to 32-bit integers with sign extensi
by derf_ 3y ago
You run into these problems on x86, also, with SIMD intrinsics like _mm_cvtepi8_epi32() (which converts four 8-bit integers to 32-bit integers with sign extension). The underlying instruction, PMOVSXBD, when used with a (4-byte) memory operand, imposes no alignment restrictions. But the intrinsic takes an __m128i value (not a pointer), and __m128i (the C type) requires 16-byte alignment.
If you try casting your pointer to (__m128i *) and dereferencing, sometimes the compiler will optimize it correctly (to PMOVSXBD with a memory operand, which really only loads 4 bytes), and sometimes (for example, when optimizations are disabled) it will emit MOVDQA, which does a 16-byte aligned load. Tools like UBSan will also (correctly) complain about it.
Eventually __mm_loadu_epi32() was added, but it was released in a broken state in the initial gcc implementation: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=99754 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=99754 so you can't rely on it. There is __mm_load_ss(), but gcc's implementation derefs a (float *), so it still requires 4-byte alignment: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=84508 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=84508
The workarounds described in the article (combined with _mm_cvtsi32_si128()) are the only reliable solution I have found. Various compilers still generate an extra MOVD instruction instead of using a memory operand, and reportedly ICC will even do an extra load into a general-purpose register before moving the value to a SIMD register: https://stackoverflow.com/questions/72837929/mm-loadu-si32-not-recognized-by-gcc-on-ubuntu https://stackoverflow.com/questions/72837929/mm-loadu-si32-n...
Those extra moves are cheap on x86, so it is not the end of the world, but the whole situation is less than ideal.
- CoastalCoder 3y agoReally interesting, thanks! I'm a little surprised about the icc thing, since it has/had a reputation for really good Intel x86 codegen.
- rwmj 3y agoParticularly insidious if you don't stick to the 16 byte stack alignment. We had one case where we calling from assembly code back to C, but the asm didn't honour the C stack alignment. The C code was trying to access some vector variable on the stack and crashed because the compiler generates code assuming %rsp is 16 byte aligned when you enter the function. The stack was only unaligned some of the time and the crash happened very distant from the wrapper, making it a pain to track down. (Asm wrapper has since been fixed).
- derf_ 3y agoIt gets worse than that. On x86-32, you are only guaranteed 8-byte stack alignment, and if you try to use __attribute__((aligned(16)) on a stack variable, gcc just silently fails to enforce the alignment. Or it did 15 years ago when x86-32 was relevant. A lot of times the code works anyway by luck. Until it doesn't. No idea if this was ever fixed.
- Am4TIfIsER0ppos 3y agoWasn't that a Windows limitation? I recall that trouble from ffmpeg having to align the stack for functions that might be called from microsoft compiled code. I recall gcc assuming alignment but no guarantee that you were actually called like that.
- ack_complete 3y agoIntel's intrinsics are just a total mess. Another example is _mm_loadl_epi64(), which loads 64-bits from memory and maps to MOVQ x,m -- but it takes an __m128i pointer. The result is that you frequently have to reinterpret cast pointers to use the intrinsics, in ways that would be blatantly broken in any other situation. ARM does a much better job of this in properly using void pointers or providing typed loads and store intrinsics. It's also fun how intrinsics are mixed between ISA extension levels with confusingly similar names that make it really easy to use them in the wrong code path. Arithmetic shift right immediate for int16 (_mm_srai_epi16) and int32 (_mm_srai_epi32) are SSE2. But _mm_srai_epi64 for int64 is AVX512. One thing that Intel does do better is that they have a standardized, OS-independent way of testing for ISA extensions through CPUID. ARM doesn't and you are at the mercy of the OS to provide you APIs to test whether, for instance, Crypto, CRC32, CAS, and UDOT instructions are available.