3 ms·
As someone who has done a lot of atomics programming: stick with sequential consistency unless you are using the variables in a very well known pattern (i.e. th
by ragnot 3y ago
As someone who has done a lot of atomics programming: stick with sequential consistency unless you are using the variables in a very well known pattern (i.e. the double locking pattern, ring buffers, etc).
Additionally, atomic doesn't mean lock free. But if they are lock free [0] you can use atomics in shared memory across a process barrier which is a lot of fun. Another good resource on this is: https://preshing.com/20120612/an-introduction-to-lock-free-programming/ https://preshing.com/20120612/an-introduction-to-lock-free-p...
[0] https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free
- colanderman 3y agoSequential consistency is always "correct" in that you can substitute it for the others and your program will be correct -- but it's generally not warranted in my experience, to the point that I find liberal use of it to be a code smell, indicative that someone may have just sprinkled them around and "hoped for the best" rather than analyzing ordering requirements of their algorithm.
- tialaramex 3y ago> stick with sequential consistency unless you are using the variables in a very well known pattern This has been popular advice, and is presumably why WG21 chose to default the ordering to sequentially consistent when standardising these types. But, if you look closer it doesn't necessarily make sense. Sequential Consistency is very expensive. Every execution context must ensure that it sees everything marked Sequentially Consistent in the same order, no matter what that order is, and so your compiler and then CPU is likely going to treat this very pessimistically. In contrast there's a good chance your high-end CPU in a server or personal computer can do Acquire-Release nice and fast if your data layout is friendly, and almost anything that's not a toy can do Relaxed [edited to correct] very fast indeed because that's easy. So if you could correctly have used Acquire-Release, but you picked (or in C++ allowed the default) Sequentially Consistent because you were scared, you are leaving orders of magnitude of performance on the table. At the other side, Sequential Consistency doesn't ensure correctness. If your algorithm wouldn't deliver what you intended under any memory ordering, Sequential Consistency can't re-order things so that it works - it's still busted although I guess debugging it might be easier now since all your execution contexts agree about the order in which things went wrong. I think the better advice is probably something like, "If you're not sure, Sequentially Consistent isn't what you wanted, what you wanted is to not use atomic memory ordering".
- ragnot 3y agoFrom my understanding, SeqCst on x86 is free since the underlying CPU has a strong memory model anyway. Only on other platforms do you get the performance benefits.
- dataflow 3y agoThat is incorrect but a common misconception (I used to believe the same). Try using it and you'll see the generated code is slower for stores on x86: https://godbolt.org/z/KrzT9bTKf https://godbolt.org/z/KrzT9bTKf I agree with GP - the commonly given "always use sequentially consistent" advice is not really good advice. It should be "use a mutex wherever possible", but once you decide that's not performant enough, you probably often do want acq/rel or relaxed. I've actually found seq_cst to be quite rare.
- gsliepen 3y agoIn addition to what dataflow already said: while something might look free because you don't need to emit any kind of special instruction to get the desired result, doesn't mean that it's free in hardware. The reason why a lot of CPU architectures have weaker memory ordering is because this is faster on a hardware level. Consider that for sequential consistency to work on a multithreaded system, caches of all cores must be kept in sync. This is hidden from you, but does result in higher latencies and more power usage.
- quietbritishjim 3y agoThe page you linked to [0] says: > The C++ standard recommends (but does not require) that lock-free atomic operations are also address-free, that is, suitable for communication between processes using shared memory. So you must check your compiler's documentation before using lock-free atomics in shared memory across a process barrier.