8 ms·
Layout of Rust's u128 and i128 changed
- ww520 3y agoThat was painful. Glad the alignment is fixed now.
- vsnf 3y agoThis is a very good article, written well for an audience that might not understand the details of memory layouts. Also, in the 0x0x11223344556677889900aabbccddeeff sample value they're using, what is this 0x0x format? Is this some some low level asm thing, or is it just a formatting typo?
- paavohtl 3y agoIt's a typo.
- Karellen 3y agoI think that the two-column compatibility table would be a lot more readable as a 2D compatibility matrix. I'm having difficulty getting a feel for what's compatible with what using the current layout.
- jupp0r 3y ago__int128 is not a C type, it's a proprietary extension of GCC and clang (among others).
- shikon7 3y agoStill it is mentioned in the ABI specification, although as an optional type.
- raverbashing 3y agoThank you compiler developers (this time) for giving resources that people use and find useful (Still weird why would LLVM not align it to 16 bytes though)
- kryptiskt 3y agoIt's a C type, it's not in the standard, but there are a lot of types that aren't in the standard. And of course, it would have been in the standard long ago if they hadn't made a terrible mistake in C99 with intmax_t[0]. [0] A Special Kind of Hell - intmax_t in C and C++: https://thephd.dev/intmax_t-hell-c++-c https://thephd.dev/intmax_t-hell-c++-c
- stephencanon 3y ago_BitInt(128) is a C(23) type and has all of the same problems plus a few others.
- tmgross 3y agoElaborating on the problems - _BitInt(128) has an alignment of 8, meaning that C's _BitInt(128) and __int128 are incompatible, similar to the Rust-C incompatibility that was just fixed :(. https://groups.google.com/g/x86-64-abi/c/-JeR9HgUU20 https://groups.google.com/g/x86-64-abi/c/-JeR9HgUU20
- stephencanon 3y agoIt's actually worse than that; from what I can piece together of the history: 1. the people at Intel who originally implemented `_BitInt()` made `_BitInt(128)` eight-byte aligned 2. the x86_64 psABI document was updated to agree with that 3. it was implemented in clang in such a way that it also made `_BitInt(128)` eight-byte aligned on arm64 ... 4. ... but the AAPCS says that it's sixteen-byte aligned 5. ... and maybe the x86_64 psABI is going to change to say that it's sixteen byte aligned after-all. So--as of a month ago when I was studying this because we're in the middle of enabling `[U]Int128` in Swift--_BitInt(128) didn't agree with __int128[_t] on either platform, and that's definitely a bug on arm64, and, although it behaves as documented on x86_64, that also might change.
- codedokode 3y agoI don't understand why align 128-bit values to 16 bytes and waste precious memory if CPUs read and process data in 8-byte chunks anyway. Or is alignment necessary for SSE insructions? But SSE doesn't work with 128-bit integers. Also, if program uses lot of memory (due to alignment) it can cause swapping and the performance will be much worse than with unaligned storage.
- riedel 3y agoAs the article explains this is purely about about compatible calling conversations between rust and C. I guess your question is also directed C compiler implementations that even overwrite LLVM defaults to achieve this alignment.
- jsheard 3y agoDoing 128bit atomics with CMPXCHG16B requires 16 byte alignment, and AFAIK that is one of the more common uses of 128bit types in practice since it's used in certain concurrency primitives to avoid the ABA problem.
- dist1ll 3y agoWouldn't it be better to expose 128-bit vector registers for this purpose? Like how the aarch64 module exposes int16x8_t. That seems much better than relying on a generic 128 bit type, because the use-case is clearly specified.
- pclmulqdq 3y agoSpecifically for CMPXCHG16B, the normal case is for this to actually be two separate numbers (usually an 8-byte pointer and an ABA counter) stored in memory as a pair. For a lot of 128-bit arithmetic, it's also better to use the integer registers to take advantage of the ADC/ADCX/ADOX and MULX instructions for basic arithmetic operations. It actually depends a lot on the operation chain.
- pclmulqdq 3y agoSSE instructions are not all that uncommon with bignums (including 128-bit types), but generally avoiding your structs crossing cache lines is very useful for performance. There are also instructions like CMPXCHG16B that need 16 bytes and are a lot worse if they cross cache lines. IMO it's actually a pretty big performance bug for 128-bit ints to not be 16-byte aligned.
- fanf2 3y agoSo much wtf in this blog post! LLVM had multiple ABI conformance bugs in its implementation of int128, but clang had workarounds for these bugs that rustc lacked, causing interop problems. I wonder why they put workarounds in clang instead of fixing LLVM?! This is a great example of one of the criticisms of LLVM, that language front ends must still implement processor-specific details themselves – iirc another example is struct layout.
- ajross 3y ago> I wonder why they put workarounds in clang instead of fixing LLVM?! Among other things, because it would probably break rust (and other projects that had come to rely on the mistaken representation). Interoperability bugs like this often end up fundamentally unfixable in practice and the community ends up having to embrace bifurcated standards.
- dylnuge 3y agoIndeed, it looks like they even tried to fix it in LLVM (back in 2017) but wound up reverting it: https://reviews.llvm.org/D28990 https://reviews.llvm.org/D28990
- fanf2 3y agoThat’s one of the links in the original article, which postdates the workaround in clang. It doesn’t explain why the incorrect ABI was previously worked around in clang instead of fixed in LLVM.
- lamontcg 3y agoAnd this is the kind of convoluted situation that happens when you have software that can never, ever break backwards compatibility. It is probably entirely necessary to do it in this particular situation, but when people argue for no software ever breaking backward compatibility, they're arguing for more and more convoluted solutions like this. There's always a cost.
- 3y ago
- ComputerGuru 3y agoI was expecting this statement to be expounded: > Unfortunately this meant some of the performance wins needed to be sacrificed to avoid an increased memory footprint.
- loeg 3y agoI think the idea is they might reduce the alignment in places to save memory.
- wongarsu 3y agoIn case anyone wants to try the same, my assumption is that they specified `#[repr(packed(8))]` on some structs to use alignments and paddings of at most 8 bytes. Using `[repr(align(8))]` on the field should also work, if you want more fine-grained control. https://doc.rust-lang.org/reference/type-layout.html#the-alignment-modifiers https://doc.rust-lang.org/reference/type-layout.html#the-ali...
- thayne 3y ago> my assumption is that they specified `#[repr(packed(8))]` on some structs to use alignments and paddings of at most 8 bytes That would mean taking a reference of any field of that struct is undefined behavior, unless the compiler does some special magic for this specific case. > Using `[repr(align(8))]` on the field I don't think you can do that directly. You would need to use a newtype that specified the alignment, and use that type for the field.
- thayne 3y agoThat would make sense, but I would like to know in which situations that happens. And can I manually reduce the alignment in a struct (say to an alignment of 8 with a struct that has other fields with alignment of 8) to reduce memory usage? Or ensure that a location uses an alignment of 16 in places I want the higher performance.
- fbdab103 3y agoWhat do you do with a 128bit integer? I already kind of consider 64 bit to be infinite. Maybe some performance optimization hack where you can use a big integer instead of a float for faster math? Skip right past 64bit unix time and give yourself a lot of breathing room?
- loeg 3y agoIt's a nice scalar representation of a UUID, for example. Also it makes a nice internal state for a fast 64-bit PRNG (e.g., any 128/64-bit LCG/MCG, such as pcg64).
- ainar-g 3y agoIPv6, too. (Unless you also need interface IDs for link-local addresses.)
- jeroenhd 3y agoI have used 128 bit (and even 256 bit) numbers during Advent of Code when I was too lazy to optimize my algorithm and found that the puzzle input didn't exceed the 64 bit space by enough to look into arbitrary sized integers or better algorithmic solutions. There are also some data formats that are 128 bit, like others mentioned.
- pclmulqdq 3y agoCryptography, UUIDs and some other hashing things, certain kinds of counters. Also, 2^64 is only about 10^19, so if you happen to have 20 exabytes that you want to byte-address, you can't do it with 64-bit pointers.
- breckognize 3y agoWhen I worked on S3, I was briefly responsible for reporting waste in the system. The basic equation was [Total Capacity of Hard Drives] - [# of bytes customers are paying for] * [Replication factor] = Waste One week as I was preparing the report, it was clear something had gone haywire. Waste was roughly equal to total capacity. So either we'd lost all of our customers overnight, or there was a bug. Turns out the legacy billing system was using a long to count # of paid bytes and this had overflowed. So it does happen.
- loeg 3y agoOne thing you can run into with these types in C/C++ is that the compiler assumes these types are aligned (unless you use some specific compiler attributes) and generates accesses that require alignment (e.g. MOVDQA) instead of ones that don't (MOVDQU). This is problematic if you have some custom allocator that (incorrectly) only provides 8-byte alignment, and cast allocated pointers to pointers to this 16-byte type (or a struct containing it). Not a Rust problem at all, just something to be careful of. I found the LLVM/Clang bugs mentioned in the article kind of fascinating. As far as I know Clang has supported these types (partially) for quite a long time, so it's interesting that these issues weren't fixed until quite recently (if at all?).
- acuozzo 3y agoAre there instances of compilers with targets having both 128-bit ints and alignof(max_align_t) != 16? I'm asking because I can't imagine writing a custom allocator against anything other than alignof(max_align_t). Unless you're solely targeting <= C99… why use anything else?
- tlb 3y agoYes: OSX on ARM64 has alignof(max_align_t)==8 and alignof(__int128)==16. And malloc only provides 8-byte alignment, so if you're using types that require 16-byte alignment, you have to call aligned_alloc. Also, the stack is only 8-byte aligned. Despite the alignof, int128 doesn't actually require 16-byte alignment on ARM64, so nothing goes wrong when you have an int128 within an 8-byte aligned struct. Some x86-64 SIMD types do require 16-byte alignment (depending on what instructions you load them with). The Eigen math library, for instance, is slightly faster when you tell it to assume everything is 16-byte aligned, but you have to do some work to guarantee that's true. As well as calling aligned_alloc, you have to avoid locals since the stack is only 8-byte aligned.
- stephencanon 3y ago> Yes: OSX on ARM64 has alignof(max_align_t)==8 and alignof(__int128)==16. And malloc only provides 8-byte alignment, so if you're using types that require 16-byte alignment, you have to call aligned_alloc. Also, the stack is only 8-byte aligned. Huh? Stack alignment is 16B ("The stack pointer on Apple platforms follows the ARM64 standard ABI and requires 16-byte alignment." https://developer.apple.com/documentation/xcode/writing-arm64-code-for-apple-platforms https://developer.apple.com/documentation/xcode/writing-arm6...) and memory returned by the system malloc is always 16B aligned (it was 16B aligned even on 32b x86 and ARM).
- maerF0x0 3y agoI had to look up what FFI meant - https://doc.rust-lang.org/nomicon/ffi.html#foreign-function-interface https://doc.rust-lang.org/nomicon/ffi.html#foreign-function-...
- ardel95 3y agoThe article didn’t mention this, but don’t u128s get mapped to SSE2 registers on most modern x86_64 processors, and not regular 64-bit ones?