3 ms·
On x86 CPUs, working with latest compilers in -O3 optimization level is becoming like working with old RISC CPUs without hardware for unaligned load/store: exce
by faragon 8y ago
On x86 CPUs, working with latest compilers in -O3 optimization level is becoming like working with old RISC CPUs without hardware for unaligned load/store: except if you write your code using memcpy/memmove for unaligned
load/store, you'll get caught by some crash. E.g. in x86-64 in -O3 GCC emits SIMD instructions that can crash your code if not properly aligned. E.g. the typical cast to 32 or 64 bit that worked in C on x86 processors, is no longer guaranteed to work:
uint8_t * p = some_address_not_aligned;
uint32_t i = * (uint32_t *)p;
In order to be safe previous code should be replaced by:
memcpy(&i, p, sizeof(uint32_t));
(the compiler is smart enough for inlining the above memcpy call, so is as fast as the previous case)
- Kenji 8y agoE.g. the typical cast to 32 or 64 bit that worked in C on x86 processors, is no longer guaranteed to work: Sorry to disappoint you, but this fragment of code was never guaranteed to work by the C standard, and x86 is a special case where it might work (but is not guaranteed to!). I hope you didn't make such type punning mistakes in the code you've written so far. A memcpy is always necessary to transform 4 bytes in an (unaligned) char array into a 32-bit integer. And yes, it all gets optimised away. At the end of the day, if you do it with memcpy, you end up with literally the same assembly as when you use a type cast, at least on GCC last time I tried. Still, it's good to not cause undefined behaviour.
- saagarjha 8y agoI don't know why this comment was killed: it's correct. The code above violates the C standard, so you're luckily that it even crashes at all rather than doing something even worse like masking to a 4-byte boundary.
- faragon 8y ago"was never guaranteed to work by the C standard" It was, on x86, before of compilers emitting opcodes not supporting unaligned memory accesses, because x86 CPUs guaranteed safe operation on unaligned memory accesses. Don't get me wrong, of course it is a good practice writing code not relying on CPU-specific characteristics. The problem is that there is a ton of software for x86 that becomes not reliable when recompiled with -O3 on x86-64 (with hard to catch situations, e.g. passing tests, but with random crashes depending on the data being handled, etc.).
- heinrich5991 8y agoWhat do you mean by "guaranteed"? Was there some specification (e.g. compiler manual or so) that said it? Or did you rely on not-specified behavior? If it's the latter, I wouldn't call that "guaranteed".
- faragon 8y agoIntel CPUs guaranteed correct operation on unaligned accesses. In fact, back in the day, transparent unaligned memory access support, and a strong memory model for SMP, were very strong selling points for Intel vs most RISC vendors. The "problem" on Intel CPUs started when SIMD instructions not supporting unaligned memory were being used for optimizing generic code. Which is a very good thing, of course. The only problem is that there is many bad quality code around that was written with the assumption of x86 CPUs being "safe" in that regard.
- obl 8y agoThat's fine if you're writing assembly. Compilers are free to use UB from the standard to optimize and absolutely do not guarantee you to emit machine loads and store naively as specified in the C source. At least for clang I'm pretty sure they decided _not_ to give knowledge of, e.g., the zero low bits of an int pointer to the optimizer since many people are relying on it. That might not be the case in the future or another compiler. Here is a thread discussing that http://lists.llvm.org/pipermail/llvm-dev/2016-January/094012.html http://lists.llvm.org/pipermail/llvm-dev/2016-January/094012...
- faragon 8y agoSure. And I love compilers generating very fast code. My point was that there is a ton of code written for x86 on assumptions that are no longer true when compiled with -O3 flags. Fortunately, open source code is less affected because usually target the generic case. However, for private code is a huge problem and risk. In my opinion, code intended to operate on x86 processors should be compiled with -O3 only in the case of high quality software, and if the quality can not be assured, it should be compiled with -O2 (at least when compiled with GCC and CLang).
- Asooka 8y agouint32_t is aligned on a 4-byte boundary, so the compiler optimisation is actually valid in this case. Reading from a misaligned pointer is not guaranteed to work.