3 ms·
x86 is byte addressable, but internally, the x86 memory bus is word addressable. So an x86 CPU does the shift/mask process you're referring to internally. Whi
by csense 4y ago
x86 is byte addressable, but internally, the x86 memory bus is word addressable. So an x86 CPU does the shift/mask process you're referring to internally. Which means it's actually slower to access (for example) a 32-bit value that is not aligned to a 4-byte boundary.
C/C++ compilers often by default add extra bytes if necessary to make sure everything's aligned. So if you have struct X { int a; char b; int c; char d; } and struct Y { int a; int b; char c; char d; } actually X takes up more memory than Y, because X needs 6 extra bytes to align the int fields to 32-bit boundaries (or 14 bytes to align to a 64-bit boundary) while Y only needs 2 bytes (or 6 bytes for 64-bit).
Meaning you can sometimes save significant amounts of memory in a C/C++ program by re-ordering struct fields [1].
[1] http://www.catb.org/esr/structure-packing/ http://www.catb.org/esr/structure-packing/
- mlyle 4y agoSure, unaligned access to memory is always expensive (on architectures that allow it at all). But I'm talking about retrieving the 9th to 16th bit of a word, which is a little different. x86 does this just fine/quickly, because bytes are addressable.
- moonchild 4y ago> internally, the x86 memory bus is word addressable The 'memory bus' is not architectural. Different microarchitectures implement things differently, but most high-performance microarchitectures these days have relatively efficient misaligned accesses.
- csense 4y ago> The 'memory bus' is not architectural Whether we classify the issue as "architectural" (whatever that means) is beside the point. Alignment has real effects on performance, and being aware of those effects is practically useful for working programmers. I do agree with you about one thing: The unaligned access penalty is probably less on the x86 CPU's of the 2020's than it was, say, on a 486 from the mid-1990's.