4 ms·
Back then, there weren't as good references for explaining atomic ordering, and the blog post had gotten long enough. Mentioning SeqCst was a bit of a cop out,
by jamesmunns 3y ago
Back then, there weren't as good references for explaining atomic ordering, and the blog post had gotten long enough. Mentioning SeqCst was a bit of a cop out, though both Andrea and I didn't end up using SeqCst past the inital impl anyway.
Today I would have just linked to https://marabos.nl/atomics/ https://marabos.nl/atomics/, Mara does a much better job of explaining atomics than I could have then or now.
- hinkley 3y agoBack then? Do you mean 2019? Or a different “then”? Because there was plenty of material in CS about this subject even in 2010. Java was wrestling with this twenty years ago, and databases long before that.
- jamesmunns 3y agoI meant 2019, and there weren't any materials that I would consider as clear and well defined as Mara's linked docs explaining the different orderings used by C, C++, and Rust (Relaxed, Release, Acquire, AcqRel, and SeqCst). I'm very sure there were discussions and teaching materials then, but none (that I was aware of) focused on Rust, and something I'd link to someone who had never heard of atomic ordering before.
- anonymous-panda 3y agoI think the hard part of it is that x86 only has one atomic ordering and none of the other modes do anything. As such, it’s really hard to build intuition about it unless you spend a lot of time writing such code on ARM which wasn’t that common in the industry and today most people use higher level abstractions. By databases, do you mean those running on DEC Alphas? Cause that was a niche system that few would have had experience with. If you meant to compare in terms if consistency semantically, sure but there’s meaningful differences between database consistency semantics of concurrent transactions and atomic ordering in a multithreaded concept. Java’s memory model “wrestling” was about defining it formally in an era of multithreading and it’s largely sequentially consistent - no weakly consistent ordering allowed. The c++ memory model was definitely the first large scale adoption of weaker consistency models I’m aware of and was done so that ARM CPUs could be properly optimized for since this was c++11 when mobile CPUs were very much front of mind. Weak consistency remains really difficult to reason about and even harder to play around with if you primarily work with x86 and there’s very little tooling around to validate that can help you get confidence about whether your code is correct. Of course, you can follow common “patterns” (eg loads are always acquire and stores are release), but fully grokking correctness and being able to play with the model in interesting ways is no small task no matter how many learning resources are out there.
- gpderetta 3y agoNit: x86 has acquire/release and seq_cst for load/stores (it technically also has relaxed, but it is not useful to map it to c++11 relaxed). What x86 lacks is weaker ordering for RMW, but there are a lot of useful lock free algorithms that are implementable just or mostly with load and stores and it can be a significant win to use non-seq-cst stores for this on x86
- haberman 3y agoIndeed there is different code generated by seq_cst for stores. Though for loads it appears to be the same: https://godbolt.org/z/WbvEcM83q https://godbolt.org/z/WbvEcM83q
- gpderetta 3y agoYes, seqcst loads map to plain loads on x86.
- loeg 3y agoRe: the godbolt example, note that release semantics are not meaningful for load operations. > If order is one of std::memory_order_release and std::memory_order_acq_rel, the behavior is undefined. https://en.cppreference.com/w/cpp/atomic/atomic/load https://en.cppreference.com/w/cpp/atomic/atomic/load
- vlovich123 3y agoI would have to imagine you mean x86-64 right? I would imagine 32bit x86 doesn’t have those instructions? I’m also kind of curious if a lot of modern code compiled to x86 would see consistency issues running on old CPUs before TSO was formalized (like a p2 multiprocessor server).
- loeg 3y ago32-bit x86 has many of the same instructions, including cmpxchg8b (in models dating to the 90s).
- 3y ago
- nyanpasu64 3y agoChapter 7 doesn't test if performing loads on a reader thread makes a writer thread any slower to perform relaxed writes. Does a concurrent reader slow down writes or not (the LMAX Disruptor relies on variables with one writer and many readers, and claims it's fast), and does it depend on the CPU's cache coherence protocol?