6 ms·
It doesn't look like the LOCK prefix applies to MOV (from a quick google)? So how does it address write buffering or OOE for stores in a TSO memory model? [edi
by johnbender 12y ago
It doesn't look like the LOCK prefix applies to MOV (from a quick google)? So how does it address write buffering or OOE for stores in a TSO memory model?
[edit] "never require the hammer of mfence for correct synchronization", maybe you're confining this to correct synch. and not recovering sequential consistency (or some other semantic property).
- mattnewport 12y agoIt's a little while since I looked closely at this and it's easy to get this stuff wrong but here's what I concluded when digging into this in the context of a codebase I was maintaining that made heavy use of lock free techniques: - You don't need mfence (or sfence / lfence) on x64 (which is actually what I care about rather than x86, though they're essentially the same in this respect) to correctly implement C++11 atomics with acquire / release semantics. You can get all the guarantees you need with locked instructions. - Correctly implemented C++11 atomics implemented with locked instructions will be as fast or faster than when implemented with explicit fence instructions. - The obvious way to implement standalone fences on x64 is with fence instructions but you rarely if ever need standalone fences. Generally you are better off using atomic operations with explicit acquire / release semantics. In the codebase I was maintaining, the pre-C++11 atomic library was based around explicit fences rather than C++11 style atomic operations with acquire / release semantics attached to the operations themselves. This code was primarily written for Gen 3 consoles (PS3 / Xbox 360) and so was optimized for PowerPC. On x64 (Gen 4 consoles!) there was measurable performance overhead due to unnecessary/redundant standalone fences. We decided it was too risky to try and rewrite everything in terms of atomic operations with acquire release semantics and remove the standalone fences in the end but it seems to me that if you want to write efficient cross architecture lock free code you want to avoid standalone fences and use a C++11 style atomics library where acquire release semantics are tied to the atomic operations themselves.
- mattnewport 12y agoTo answer the first part of your question, a locked xchg is the usual way to implement a sequentially consistent store on x86 without an explicit fence. For sequentially consistent loads, plain mov suffices, providing it synchronizes with a locked store. http://stackoverflow.com/questions/4972106/sequentially-consistent-atomic-load-on-x86 http://stackoverflow.com/questions/4972106/sequentially-cons...