3 ms·
Let’s be clear here, while you are 100% right that armv8 atomics kinda suck by default, neither compare and swap nor load linked store conditional scale but som
by trws 3y ago
Let’s be clear here, while you are 100% right that armv8 atomics kinda suck by default, neither compare and swap nor load linked store conditional scale but some atomics can scale if implemented and used appropriately. As parent points out, an atomic increment can scale, we proved they could scale to the performance of a load in the 80s for goodness sake. The fact that arm, ppc, and some others tend to implement these in the absolute worst way possible for performance doesn’t mean atomics can’t scale.
- ghusbands 3y agoYou're not being clear - are you claiming that they do scale well on ARM/PPC, despite the "absolute worst" implementation or that they don't?
- bonzini 3y agoHe means that locks anyway have the same problems as atomics on platforms with load locked/store conditional. Therefore yes, that's a problem of the platform, but even on arm/ppc atomics scale better than locks. Which is true, but eliminating sharing works even better if possible as proved by the article.
- moonchild 3y ago> an atomic increment can scale, we proved they could scale to the performance of a load in the 80s Interesting--can you link a reference for this?
- trws 3y agoThis is the citation I most often use from that time, though the primary source is probably another reference down the line: https://dl.acm.org/doi/10.1145/69624.357206 https://dl.acm.org/doi/10.1145/69624.357206 The short version is that if atomics are implemented as part of the memory network, common cache, or memory controller, then atomics of the form “fetch-and-X” can be implemented in roughly equivalent complexity to a load of the current value (plus an instruction for the op, give it take) with the cost only scaling past that as op queues or other implementation-specific limits fire. It’s the infinite consensus ops that just can’t scale no matter what you do. The coherence and memory model matter a lot too of course, which is part of why x86 tends to be slow for atomics, while arm and ppc (with fetch-and-X extensions) or GPUs tend to do much better.
- temac 3y agoI'm not sure about how atomics having the performance of loads mean they will scale (and to be honest i doubt they could have the perf of loads on modern architecture, otherwise why would e.g. Intel not implement them to be faster - but lets pretend it is possible) The fine article shows that a single lock xadd can destroy perfs on some x86 systems and explain that it is due to cache line bouncing. You would get the same effect with loads: if the loaded data is RO or mostly RO it will of course scale fine. It won't scale as soon as it starts bouncing too much.
- trws 3y agoThen why keep bouncing it? Leaving it managed by a known single user in a cache architecture like on x86 means that latency goes up, but the overall throughput of the operation goes up drastically. That’s why flat combining data structures are so popular there despite their absolute maximum throughput being bounded by sequential performance. Also, FWIW, intel largely does implement operations closer to that way on single socket parts, if you want to see it for real look at on-device atomics on a GPU. Ironically an average laptop chip handles atomics much faster than most servers as a result.