4 ms·
Atomics scale very well if you are reading often and writing rarely.
by eldenring 3y ago
Atomics scale very well if you are reading often and writing rarely.
- anaisbetts 3y agoYep, uncontended atomics are quite fast. When they're contended is when things start to slow down like OP has seen
- klabb3 3y agoSorry to nit, but this is important. Parent was saying that reads are cheap, which is true. Writes can be expensive even if uncontended, because they invalidate cache lines. I guess you could say they contend with unrelated data but that would stretch the definition a bit. So what does this mean in practice? In my view, the way to think about it is that atomic writes have non-local side effects. But since atomics are necessary for synchronization, and involves both reads and writes, we should compartmentalize and minimize synchronization as much as possible, to avoid these gnarly issues creeping up and tanking real world performance. Arc<T> (and it’s relatives in other languages) constitute textbook violations of this rule. In Rust they are everywhere in non-trivial code, including in the async runtimes themselves. Of course, they also violate (or evade if you’re generous) ownership principles of idiomatic Rust, (or “hello world-Rust”, if you will). I think we need to take a hard look as an industry at ref counting as a silver bullet escape hatch to shared data.
- loeg 3y ago> Writes can be expensive even if uncontended, because they invalidate cache lines. This isn't expensive if cache lines are uncontended, though. > I guess you could say they contend with unrelated data but that would stretch the definition a bit. I think you might be talking about "false sharing." This is real contention on the cache line due to co-location of apparently unrelated variables. > Arc<T> (and it’s relatives in other languages) constitute textbook violations of this rule. Definitely! > In Rust they are everywhere in non-trivial code Ehh.. only the hot ones matter. Most are not actually contended much, and the article's solution (unshared clone) is a very reasonable approach to scale these without an API change.
- klabb3 3y ago> I think you might be talking about "false sharing." This is real contention on the cache line due to co-location of apparently unrelated variables. You’re right. And cache lines are quite small, so this is probably less common. Yet, it’s another potential source of perf regressions in concurrent code, as if it wasn’t incredibly complex already. > Ehh.. only the hot ones matter. Well.. first atomics have even more non-local effects, such as barriers on instruction reordering. So Arcs that are cloned willy nilly can still be significant, with no contention. But let’s ignore that and focus on the contended case: when you hear “uncontended X are basically free” it (subjectively, imo) downplays the issue, like contention is some special case that you can compartmentalize and only worry about when you consciously decide to write contended code. The blog post demonstrates exactly how this is so easy for contention to creep in, that you have to be superhuman levels of vigilant and paranoid to spot these issues upfront. Extremely easy to miss in eg code review. I think both compile- and runtime tooling could help at least partly here. I’d also give rust some credit for having explicit clone instead of hiding it.
- anaisbetts 3y agoIt's a good point, it's easy to cause contention and not realize it because of cache lines
- gpderetta 3y agoExactly. Atomics are a red herring. The single writer principle should be the fundamental guideline.
- magicalhippo 3y agoI was optimizing some code where, for reasons, many threads had to aggregate data in a shared data structure. Each thread would have a small buffer, and when full it would acquire a lock and aggregate. While fooling around I had the idea to exploit that not all atomic operations are equal. So I added an additional "contention" flag. When a thread wanted to aggregate it would do an atomic read of the flag, if it was set it would bail[1] and continue to accumulate to the local buffer. Once done aggregating the flag would be reset after unlocking. Effectively this was adding a single-iteration spinlock before the "heavy" lock, but even using CriticalSections for the lock on Windows (which does spin before acquiring a mutex lock) it resulted in clear improvements, especially when running on machines with more than 8 cores. [1]: It would bail unless the local buffer grew too large, so slight memory vs perf tradeoff there.
- loeg 3y agoThis is more or less equivalent to a try_lock() operation, yeah? And continuing to collect to the local buffer if try_lock fails, up to some limit.
- magicalhippo 3y agoYeah I suppose. One limitation we had was cross-compiling on Windows, Linux and OSX so available locking primitivea differed, so was easier to just do it via atomics.