4 ms·
The intro of this article repeats the common assertion that > ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict,
by pdw 15d ago
The intro of this article repeats the common assertion that
> ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization
but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit, https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-cpus/#320ecacb-31dd-4a74-bd34-de2dbf46b1e0 https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-...
- Aissen 15d agoA recent study seemed to support Fabian's well written article: https://dl.acm.org/doi/epdf/10.1145/3779212.3790129 https://dl.acm.org/doi/epdf/10.1145/3779212.3790129
- wat10000 15d agoIs this just a question of what one considers to be “significant”? I’d consider 3% to be significant but the authors apparently don’t.
- deater 15d agowell it depends on your error bars. On many modern system you can have +/- 5% variation or more run to run just due to the non-determinism present in modern CPU architectures and operating systems (even things like room temperature, time of day, the number of environment variables, etc, can affect this). While you could maybe run a set of careful experiments to characterize and remove this, in my experience most researchers don't bother. So something as small as 3% would need a lot of convincing to me to make the argument that it is significant.
- wat10000 15d agoMultiple separate questions here. First, is a measured improvement actually real or just an artifact of noise? Second, if it is real, is 3% anywhere close to the true value? Third, if 3% is real, is it an important difference? I'm just commenting on the third one. If 3% is real, it's important. I'm inclined to believe there's a real improvement. They made a lot of different measurements. If the measured improvement was a result of noise, you'd expect a lot of variation an a lot of measurements where TSA was actually faster, and then 3% was the average of that variation. There was a lot of variation (expected, because they were measuring different things) but nearly all of them had TSO being either neutral or slower. Looking at their benchmark graphs, I see two (out of dozens) where TSO was faster. As far as being close to the true value, these results suggest there is no single true value, as it depends on the workload. No surprise there.
- kccqzy 15d agoIt’s a question of how much resources to allocate to the hardware team, and how much resources to be distributed diffusely to the software engineers but especially to the compiler team. Even your linked paper contends that the actual observed slowdown is as much as 22% in the Geekbench example, but the thesis is that the slowdown is not inherent to TSO, but merely to the specific hardware implementation. Is it worthwhile for a company to optimize its TSO to chase the final gains, or is it better not to have this feature in the first place and just change the compiler? Indeed my instinct is that it is better to do this in software, where the programmer clearly communicates which stores are ordered, and which may happen in arbitrary order.
- cwzwarich 15d ago[Disclaimer: I wrote Rosetta 2 and determined the spec for Apple's TSO mode, so I am obviously biased.] Giesen's article comes off as well-meaning cope from an x86 fan. A relaxed memory model really does give you some performance. Another memory model flaw here in x86 is more architectural, which is that every instruction with the LOCK prefix is essentially a full barrier (of course, x86 could have provided different instructions while still being under TSO). In programs that make heavy usage of atomic reference counting, this actually helps quite a bit. I would probably put that performance benefit in the single digit percentage range like my sibling comment, which may not seem like much to a SW engineer but is actually pretty serious in CPU microarchitecture. It also helps to be stacked with other architectural advantages over x86, e.g. fixed-length instructions, 32 GPRs (which Intel copied in APX), LDP/STP (which Intel also copied in APX), etc. One of the old arguments from TSO enjoyers was that TSO helps avoid concurrency bugs that people would accidentally introduce, but this was before the C++ memory model propagated throughout the programming world. Nowadays, I think people generally conceptualize memory consistency in terms of acquire/release anyways, so why not use a CPU architecture that uses the same model?
- spijdar 15d agoDo you think there would be any worthwhile gains from relaxing address-dependent load ordering, like on Alpha/AXP? Or was that just a lot of extra pain for little reward?
- cwzwarich 15d agoFun fact: ARM actually has relaxed address dependencies for non-temporal loads, although I don't know if too many implementations of ARM take advantage of this relaxation. I think it is an interesting question. During the Alpha's lifetime as a non-hobbyist architecture, this decision was pretty much universally derided, but this was in the prehistoric eras of concurrent memory models, where people were just trying their best with a mix of C code, intrinsics, uses of `volatile` sprinkled around to hopefully disable optimizations, and inline assembly. When the C++11 memory model came around, they tried to integrate dependency ordering with `memory_order_consume`, and this famously failed, along with every attempt to fix it. I believe the plan is now for C++ (and later C?) to add special-cased RCU primitives. The relaxation makes sense in the abstract. As evidenced from the `memory_order_consume` saga, compiler optimizations regularly violate dependency ordering anyways, so in the strictest sense you can't really rely on it. Of course, that doesn't stop people from "knowing" what their compilers will do in such a situation, but that strategy has become a worse one as the years have passed. I would feel better about the whole situation if there was a good greenfield design for a low-level PL that incorporates explicit dependency ordering. I think it's less clear where the potential HW benefit is in a contemporary CPU. The obvious answer is value prediction, but any CPU performing value prediction has to deal with so many other microarchitectural conditions that can invalidate its speculation decisions that it's not clear this minor one is a huge burden. Of course, people who love TSO might make the same argument, i.e. that it's not a huge burden to snoop cache traffic and invalidate loads (although this is only the load half of TSO, not the store half). I know an Alpha architect who argued that Alpha was right for this decision. I even knew a Transmeta architect who argued for implementing sequential consistency in HW (as Transmeta and derived CPUs actually did). In practice, microarchitectural structures have capacity/throughput limitations, and there are implementations and complexities that only come up in a real design, so everything needs to be evaluated in the context of a real project. I personally think the sweet spot falls to the weaker side of TSO, which also happens to be near the memory consistency model of the low-level languages we're using anyways.
- dmitrygr 15d ago> but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit, "ELI5" oversimplified explanation why this is insane and can be no other way: You write A (not in cache), you write B, you write C, ... you write Y, you read Z (in L1d). In ARM, when A misses L1d, L2, and L3, you can issue writes to B..Y, and the read of Z (which hits in 2 cycles) and go on your merry way, your pipeline sure of the value of Z. With TSO you CANNOT allow the writes to B..Y to be seen by anyone before A, because everyone expects to only see them after A, so you either buffer them or, eventually, sit on your proverbial ass and wait for A to drain. You might have the buffers to put off the consequences for a while, but that is explicitly a cost ARM lacks, AND(this one hurts the most) you can't simply let the load of Z become an architecturally committed observation before A is done. You can speculate it, but you have to preserve the TSO ordering constraints. Eventually you'll run out of write-buffer slots or other resources, or find that your guess of Z's value was ... wrong. This is an example of how TSO loses perf compared to relaxed. The list of how it gains perf compared to relaxed is shorter: { } So, in some cases TSO is the same perf as relaxed, in no cases it is faster, in some cases it is slower. So it must be slower overall, since "overall" is a weighted mix of those. "By how much" is a question of detail and workload. However, unless your workload miraculously never suffers a cache miss while a later instruction hits, TSO must be slower. :)