5 ms·
I wonder how it handles the stricter memory ordering from x86? For example, this code: void store_value_and_unlock(int *p, int *l, int v) { *p = v;
by andoma 6y ago
I wonder how it handles the stricter memory ordering from x86? For example, this code:
void store_value_and_unlock(int *p, int *l, int v)
{
*p = v;
__sync_lock_release(l, 0);
}
on x86_64 compiles to:
mov dword ptr [rdi], edx
mov dword ptr [rsi], 0
ret
vs on arm64:
str w2, [x0]
stlr wzr, [x1]
ret
Notice how the second store on ARM is a store with release semantics to ensure correct memory ordering as it was intended in the original C code. This information is lost as it's not needed on x86 which guarantees (among other things) that stores are visible in-order from all other threads.
- monocasa 6y agoThat's the big piece I've been wondering about too. Three options as I see it (none of them great): 1) Pin all threads in an x86 process to a single core. You don't have memory model concerns on a single core. 2) Don't do anything? Just rely on apps to use the system provided mutex libraries, and they just break if they try to roll their own concurrency? Seems like exactly the applications you care about (games, pro apps), would be the ones most likely to break. 3) Some stricter memory model in hardware? Seems like that'd go against most of the stated reason for switching to ARM in the first place.
- andoma 6y ago> 3) Some stricter memory model in hardware? Seems like that'd go against most of the stated reason for switching to ARM in the first place. I would assume that a more strict memory model would be enabled only for processes that needs it (ie, Rosetta translated ones). So a cpu-flag is set/cleared when entering/exiting user mode for those processes. Does this require a separate/special cache coherency protocol? A complete L1d flush when entering/leaving these processes (across all CPUs)? Not and expert in this field and it feels complicated for sure. Is it worth it for just emulating "legacy" applications during a transitional period? Perhaps Apple can pull it off though.
- monocasa 6y ago> I would assume that a more strict memory model would be enabled only for processes that needs it (ie, Rosetta translated ones). So a cpu-flag is set/cleared when entering/exiting user mode for those processes. Does this require a separate/special cache coherency protocol? The benefits you'd get from going to a weaker memory model are by not having that extra coherency in the critical path in the first place. Adding extra muxes in front of it to make it optional would be worse than just having it on all the time. > A complete L1d flush when entering/leaving these processes (across all CPUs)? That wouldn't help because two threads could be running at the same time on different cores against their respective L1 and store buffers.
- andoma 6y ago> The benefits you'd get from going to a weaker memory model are by not having that extra coherency in the critical path in the first place. Adding extra muxes in front of it to make it optional would be worse than just having it on all the time. Indeed true, good point. > That wouldn't help because two threads could be running at the same time on different cores against their respective L1 and store buffers. Of course, this was related to the cost of switching coherency protocol during context switch. But as you say the overhead of just making it switchable is prohibitive in itself.
- phire 6y agoIt can be enabled per instruction. Atomic instructions (and ARMv8.1 added a bunch of new atomic read-modify-write instructions that line up nicely with x86) use the new cache coherency protocol, while the older non-atomic instructions keep the relaxed memory model. Though, I'm not sure if it's worth it to keep two concurrency protocols around. I wouldn't be surprised if the non-atomic instructions get an undocumented improvement to their memory model.
- BeeOnRope 6y agoThere's another option: 4) Translate x86 loads and stores to acquire load and release stores, to align with the x86 semantics. These already exist in the ARM ISA, so it's not much of a stretch at all. This is the one I'm betting on.
- monocasa 6y agoExpect you'd need to do that with every single store since the x86 instruction stream doesn't have those semantics embedded in it. That'd kill your memory perf by at least an order of magnitude, and kill perf for other cores as well. It'd be cheaper to just say "you only get one core in x86 mode". Essentially you'd be marking every store as an L1 and store buffer flush and only operating out of L2.
- my123 6y agoWhat I can say is that they aren't pinning to only a single core, so the answer is elsewhere.
- phire 6y agoARMv8.1 adds a bunch of improved atomic instructions that basically implement the same functionality as x86. Because x86 has atomic read-modify-write instructions by default; You need emulate those too. The ARMv8.1 extensions look like they have been explicitly designed to allow emulation of x86. The implication is that implementations of ARMv8.1 can (and perhaps should) implement cache/coherency subsystems with high preformance atomic operations. And I'm willing to bet Apple has made sure their implementation is good at atomics.
- monocasa 6y agoSo the ARMV8.1 extensions aren't for emulating x86, they're a reflection of how concurrency hardware has changed over the years. It used to be that (and you can see this in the RISC ISAs from the 80s/90s, but AIUI this is what happend in x86 microcode as well) * the CPU would read a line and lock it for modifications in cache. * the CPU core would do the modification * the CPU would write the new value down to L2 with an unlock, and L2 would now allow this cache line to be accessed So you're looking at ~15 cycles of contention from L2 through the CPU core and back down to L2 of a lock for that line. If this all looks really close to a subset of hardware transnational memory when you realize that the store can fail and the CPU has to do it again if some resource limit exceeded, you're not alone. Then somebody figured out that you can just stick an ALU directly in L2, and send the ALU ops down to it in the memory requests instead of locking lines. That reduces the contention time to around a couple cycles. These ALU ops can also be easily included in the coherency protocol, allowing remote NUMA RMW atomics without thrashing cache lines like you'd normally need to. This is why you see both in the RISC-V A extension as well. The underlying hardware has implementations in both models. I've heard rumors that earlier ARM tried to do macro op fusion to build atomic RMWs out of common sequences, but there were enough versions of that in the wild that it didn't give the benefits they were expecting. However, all that being said, atomics are orthogonal to what I'm talking about. It's x86's TSO model, and how it _lacks_ barriers in places that ARM requires them, with no context just from the instruction stream about where they're necessary that's the problem here, not emulating the explicit atomic sequences.
- AndrewStephens 6y agoElsewhere in the documentation[0] Apple explicitly calls out that code that relies on x86 memory ordering will need to be modified to contain explicit barriers. All sensible code will do this already. [0] https://developer.apple.com/documentation/apple_silicon/addressing_architectural_differences_in_your_macos_code https://developer.apple.com/documentation/apple_silicon/addr...
- gpderetta 6y agoSure, for source translation that's sorta fine (although I wouldn't want to be the one debugging it). The issue is binary translation, there is really no such a thing as a acquire or release barrier on x86.