4 ms·
I don't understand why align 128-bit values to 16 bytes and waste precious memory if CPUs read and process data in 8-byte chunks anyway. Or is alignment necessa
by codedokode 3y ago
I don't understand why align 128-bit values to 16 bytes and waste precious memory if CPUs read and process data in 8-byte chunks anyway. Or is alignment necessary for SSE insructions? But SSE doesn't work with 128-bit integers.
Also, if program uses lot of memory (due to alignment) it can cause swapping and the performance will be much worse than with unaligned storage.
- riedel 3y agoAs the article explains this is purely about about compatible calling conversations between rust and C. I guess your question is also directed C compiler implementations that even overwrite LLVM defaults to achieve this alignment.
- jsheard 3y agoDoing 128bit atomics with CMPXCHG16B requires 16 byte alignment, and AFAIK that is one of the more common uses of 128bit types in practice since it's used in certain concurrency primitives to avoid the ABA problem.
- dist1ll 3y agoWouldn't it be better to expose 128-bit vector registers for this purpose? Like how the aarch64 module exposes int16x8_t. That seems much better than relying on a generic 128 bit type, because the use-case is clearly specified.
- pclmulqdq 3y agoSpecifically for CMPXCHG16B, the normal case is for this to actually be two separate numbers (usually an 8-byte pointer and an ABA counter) stored in memory as a pair. For a lot of 128-bit arithmetic, it's also better to use the integer registers to take advantage of the ADC/ADCX/ADOX and MULX instructions for basic arithmetic operations. It actually depends a lot on the operation chain.
- pclmulqdq 3y agoSSE instructions are not all that uncommon with bignums (including 128-bit types), but generally avoiding your structs crossing cache lines is very useful for performance. There are also instructions like CMPXCHG16B that need 16 bytes and are a lot worse if they cross cache lines. IMO it's actually a pretty big performance bug for 128-bit ints to not be 16-byte aligned.
- gpderetta 3y ago"Worse" is a slight understatement. When not disabled by the OS, cacheline crossing RMW are 4-5 orders of magnitude slower than normal and affect all CPUs in a system.
- PartiallyTyped 3y agoIn general, Rust fields are padded such that they aligned to a multiple of their size. Rustc does not offer guarantees on the ordering of the fields. This holds for all types, not just 128bit values. This allows rust programs to actually use less memory than C programs.
- kibwen 3y ago> Rustc does not offer guarantees on the ordering of the fields. It's possible to guarantee the layout of a struct by opting into a specific representation, such as the `repr(C)` shown in the OP.
- PartiallyTyped 3y agoThat's a special case and sits at the boundaries / bridge of languages. > There is no indirection for these types; all data is stored within the struct, as you would expect in C. However with the exception of arrays (which are densely packed and in-order), the layout of data is not specified by default. struct A { a: i32, b: u64, } struct B { a: i32, b: u64, } > Rust does guarantee that two instances of A have their data laid out in exactly the same way. However Rust does not currently guarantee that an instance of A has the same field ordering or padding as an instance of B. https://doc.rust-lang.org/nomicon/repr-rust.html https://doc.rust-lang.org/nomicon/repr-rust.html
- codeflo 3y agoIt’s not the default, but it’s just a language feature. It’s also not just for interfacing with C. For example, you also use repr(C) for pointer tricks even if you never leave Rust. Lots of places in the standard library use it for that reason.
- PartiallyTyped 3y agoI don't understand the reaction. > By default, composite structures have an alignment equal to the maximum of their fields' alignments. Rust will consequently insert padding where necessary to ensure that all fields are properly aligned and that the overall type's size is a multiple of its alignment. [...] > There is no indirection for these types; all data is stored within the struct, as you would expect in C. However with the exception of arrays (which are densely packed and in-order), the layout of data is not specified by default. struct A { a: i32, b: u64, } struct B { a: i32, b: u64, } > Rust does guarantee that two instances of A have their data laid out in exactly the same way. However Rust does not currently guarantee that an instance of A has the same field ordering or padding as an instance of B. Here is my source: https://doc.rust-lang.org/nomicon/repr-rust.html https://doc.rust-lang.org/nomicon/repr-rust.html
- dragontamer 3y agoThe CPU works on 64-bits. But the memory works on 512-bit / 64 byte cache lengths or DDR4 bursts. An unaligned access could be across 2 cache lines or 2 DDR Bursts. Or across a page table (I think 4096-bytes??) Or other higher level of organization requiring multiple accesses under the hood. X86 does the easy thing and takes two or more clock ticks to read. But aligned accesses remain faster, likely for these low level groupings of bytes. Some architectures straight up do not allow unaligned reads and force the programmer to read two registers and then extract the unaligned data. ------- In particular, reading or writing across two cache lines can be disastrous loss of efficiencies. The memory system needs to lock / MESI indicate both cache blocks. If one (or the other) memory location is still locked by another thread / CPU core, you'll get false sharing. As such, a common practice is to 64-byte / 512-bit align your memory accesses in heavily multithreaded code.
- kzrdude 3y agoAnd sometimes 128-byte because the some part of the cache system is reading two cache lines at a time in contemporary x86-64 if I understand correctly.
- loeg 3y agoIt was a really big problem for early x86-64 CPUs, I believe. My hearsay recollection is that the cache prefetcher would pretty much always load pairs of adjacent cache lines.
- dist1ll 3y agoNote: unaligned doesn't always mean cache-line splitting. Unaligned 64-bit loads & stores within a cache line incur no performance penalty on modern Intel architectures IIRC. Also, reading a cache-line in shared MESI state should not cause false sharing degradation, only writing.
- pclmulqdq 3y agoUnaligned doesn't always mean cache-line splitting, but aligned means never splitting cache lines. Also, you're correct that aside from caching effects, unaligned loads and stores generally carry no penalty on x86. This is one of the smarter (IMO) things that Intel/AMD have done to keep x86's market share.
- loeg 3y agoThe "A" in MOVDQA stands for "aligned," and requires 16-byte alignment. The corresponding MOVDQU does not require alignment but is marginally slower, at least on older CPUs.
- jsheard 3y agoOn modern CPUs MOVDQU is still slower, but only if the data is unaligned and straddles two cachelines, if the data is properly aligned than MOVDQU and MOVDQA perform identically nowadays. It's still important to align data but it doesn't matter so much whether you use the aligned instructions. I suppose using the aligned instructions still gives you a free assertion that data you expect to be aligned is actually aligned, rather than silently running slower if it's not. IIRC at least one compiler (maybe ICC?) no longer bothers to emit aligned instructions at all, even if it knows the data should be aligned, because they found it to be more trouble than it's worth on modern hardware.
- pclmulqdq 3y agoIIRC MOV*A and MOV*U differ pretty significantly in how the secret stuff in the cache prefetching system treats them. The assertion is kind of nice, but there is a difference.