8 ms·
Ways to break your systems code using volatile (2010)
- nullwasamistake 7y agoIronically volatile is just as bad in Java for different reasons. Frequently used for "lock free" synchronization, its usually actually worse than using locks because it can't be cached between cores. The variable is always loaded from main memory, which is usually much worse than holding a lock mutex in registers.
- the8472 7y agoThe standard pattern for working with atomics in java (volatiles are of limited use without atomic field updaters or varhandles) is to read it into a local variable, operate on that and only write it back to the volatile once you're done. That has many benefits, among them the ability to store its value in registers.
- nullwasamistake 7y agoFor primitives Java uses special CPU instructions. In the Atomic* package. It's not recommended for plain objects.
- chrisseaton 7y ago> Frequently used for "lock free" synchronization, its usually actually worse than using locks If you want lock-free what do you suggest we use instead of volatile? > which is usually much worse than holding a lock mutex in registers How can you hold a mutex in a register? That doesn't make any sense.
- nullwasamistake 7y agoYou can use a lock but held in a regular CPU register. They're just regular variables for the most part. Lock free in Java is usually worse that what the JVM can pull off with lock elison
- chrisseaton 7y ago> You can use a lock but held in a regular CPU register. They're just regular variables for the most part. I don't understand this. If your lock variable is in a CPU register how do other CPUs acquire the lock? > Lock free in Java is usually worse that what the JVM can pull off with lock elision I don't understand this either. Java's lock elision is only going to make a concurrent object 'lock-free' in the case where the object does not escape the compilation unit. In which case again how would another thread use it? Java will also combine adjacent critical sections created by monitors even if they escape, but it won't make them lock-free in that case.
- nullwasamistake 7y agoJVM is very smart about locking. My knowledge is limited but this is a great article. https://shipilev.net/jvm/anatomy-quarks/19-lock-elision/ https://shipilev.net/jvm/anatomy-quarks/19-lock-elision/
- chrisseaton 7y agoI work with the JVM at Oracle and I’ve given talks about the lock elision algorithm. It doesn’t do what you think it does and what it does do is not related to lock-free like you think it is.
- alain94040 7y agoYou really should just use volatile for device drivers when accessing IO space with side-effects. Do not use volatile to build your own synchronization primitives.
- vardump 7y agoDon't forget about memory barriers. Otherwise your driver will fail on other CPU architectures. Just because it works on x86, doesn't mean it works on ARM, MIPS, POWER or RISC-V. CPUs other than x86 can reorder stores with other stores and loads with other loads. It can cause the CPU to do the store that starts DMA before the stores that set up length and address are done! Or just use C11 or C++11 memory model. Although those are still not available in too many cases, curse of having to use an ancient compiler...
- burfog 7y agoEven on x86, even with the C11 memory model, you can still get burned by transaction reordering as the MMIO passes through bridges. Plain old PCI can do this.
- oconnor663 7y agoI thought x86 wasn't allowed to do write-write reordering? Does that rule not apply to peripherals? Is an `mfence` guaranteed to fix it, or are there just no rules at all at that point?
- burfog 7y agoPCI isn't x86, and x86 isn't PCI. PCI itself, for example in a PCI-to-PCI bridge chip, can do store buffering and read prefetches. The PCI specification lays out what you must to to suppress this. There are at least 5 different sets of rules for ordering on x86, due to memory types. It's in the Intel manual, along with a table that shows how they interact with each other.
- deleted 7y ago[deleted]
- raphlinus 7y agoThis should say 2010. I believe much of it is out of date, as C11 does have a memory model, and does provide both atomics and barriers. Many, if not most, uses of volatile should probably be replaced by atomics. https://en.cppreference.com/w/c/atomic https://en.cppreference.com/w/c/atomic
- oconnor663 7y agoHerb Sutter's (three hour long!) Atomic Weapons talk goes over the modern meaning of volatile towards the end: https://youtu.be/KeLBd2EJLOU?t=5299 https://youtu.be/KeLBd2EJLOU?t=5299
- User23 7y agoYou're not wrong, but C is a language where plenty of apps are still using older versions, so while as per site etiquette this should have the date, it's still interesting, especially for the embedded space.
- sk0g 7y agoThere's using an outdated version, and then there's C. The more commonly used language specs, IME, are C99 and C89! That's 20 and 30 years of 'stability'.
- icedchai 7y agoMany projects are stuck in C99, or even C89...
- flafla2 7y agoEdit: Looks like the slides had an inaccuracy (see replies). Huh, looks like I learned something today :) I think a good way of summarizing volatile is this slide from my parallel architectures class [1]: > Class exercise: describe everything that might occur during the > execution of this statement > volatile int x = 10 > > 1. Write to memory > > Now describe everything that might occur during the execution of > this statement > int x = 10 > > 1. Virtual address to physical address conversion (TLB lookup) > 2. TLB miss > 3. TLB update (might involve OS) > 4. OS may need to swap in page to get the appropriate page > table (load from disk to physical address) > 5. Cache lookup (tag check) > 6. Determine line not in cache (need to generate BusRdX) > 7. Arbitrate for bus > 8. Win bus, place address, command on bus > 9. All caches perform snoop (e.g., invalidate their local > copies of the relevant line) > 10. Another cache or memory decides it must respond (let’s > assume it’s memory) > 11. Memory request sent to memory controller > 12. Memory controller is itself a scheduler > 13. Memory controller checks active row in DRAM row buffer. > (May > need to activate new DRAM row. Let’s assume it does.) > 14. DRAM reads values into row buffer > 15. Memory arbitrates for data bus > 16. Memory wins bus > 17. Memory puts data on bus > 18. Requesting cache grabs data, updates cache line and tags, > moves line into exclusive state > 19. Processor is notified data exists > 20. Instruction proceeds > * This list is certainly not complete, it’s just > what I came up with off the top of my head. It's also worth mentioning that this assumes a uniprocessor model, so out-of-order execution is still possible which leads to complications in any sort of multithreaded or networked system (See #5, 6, 7, 8 in the OP article). I think a lot of the confusion stems from the illusion that a uniprocessor + in-order execution model implies to programmers who have never dealt with system-level code. I think in the future, performant software will require a bit more understanding of the underlying hardware on the part of your average software developer -- especially when you care about any sort of parallelism. It doesn't help that almost all common CS curriculum ignores parallelism until the 3rd year or more. [1] http://www.cs.cmu.edu/~418/lectures/12_snoopimpl.pdf http://www.cs.cmu.edu/~418/lectures/12_snoopimpl.pdf - the last 2 slides
- keldaris 7y agoI don't understand these slides. The volatile keyword does not magically bypass the mechanism by which modern CPUs write to main memory. Am I missing something, or are they somehow meant to be ironic?
- loeg 7y agovolatile should only be used for accessing MMIO registers in device drivers; that's it.
- ychen306 7y agoThis is probably not useful for production, but volatile is a great way to see what kind of code compiler generates in a realistic setting. For example, if you want to see how compiler optimizes a code snippet and the code depends on a constant that you don't want to get constant folded away.
- ridiculous_fish 7y agoThere's definitely more uses. For example, shared memory between processes: you should mark it volatile. C++ atomics are no good here, because they are not guaranteed to be lock free or address free.
- ychen306 7y agomaybe use fences?
- loeg 7y agoShared memory is way outside the scope of standard C or C++. It's implementation-defined. It's inconsistent to insist on the weakest definition of atomics allowed by the C/C++ standard(s) and simultaneous invoke one of the weirdest implementation-defined mechanisms defined by POSIX. If your implementation provides shared memory of some kind, it's up to your implementation to define some sort of reasonable semantics. In POSIX' case, it's up to POSIX operating systems to define reasonable semantics on the memory, using constructs like PTHREAD_PROCESS_SHARED and "robust" pthread mutexes.
- vardump 7y ago> C++ atomics are no good here, because they are not guaranteed to be lock free or address free. That's not right; you can still use std::memory_order to get the memory barriers generated that are required. These are going to obviously be lock free, they deal with memory ordering—what you tried to deal with volatile, but in general case. See: https://en.cppreference.com/w/cpp/atomic/atomic/store https://en.cppreference.com/w/cpp/atomic/atomic/store Effectively std::atomic stores and loads generate volatile accesses plus the required memory barriers to get the desired behavior.
- tus87 7y agoErr...volatile just tells the compiler not to cache the value in a register, that's it. If you don't understand volatile you really, really are not the kind of programmer who should even think about using it.
- aidenn0 7y agoThe very first example in TFA shows the compiler doing more than this.
- empiricus 7y agoThis is my understanding of volatile as well: volatile just forces read/write to memory. What a read/write to memory entails is a different story. What happens with no volatile is again another story. If my understanding is wrong, someone please enlighten me.
- tus87 7y agoYou are correct. The author is a delusional narcissistic idiot who believes in a C "memory model" or "virtual machine". It bullcrap. C is just high level assembler, that makes certain optimizations based on certain assumptions. Lots of stuff can break at -O3 at higher. One usually fairly reasonable assumption is that values will not change between reads and writes. Sometimes they do, hence volatile. When you write actual assembly this all becomes irrelevant as you can choose where you store variables and how you access them and when you should consider them "stale". And when you understand code at this low level, C and it's volatile becomes rather plain and boring. There is a LOT of this kind of blogspam on the internet where fools obtain a semi-correct understanding of some low level concepts, then write these long essays about this "dark magic" and making themselves out to be these grandiose elite greybeards hack0rs. And guess where it rises to the front page to the adoration of semi-sentient wannabe dreamer engineers who coo and ga over how smart he is? You guessed it.
- moefh 7y ago"Forcing read/write to memory" is very different from "not caching the value in a register". Optimizations can involve not just caching values in registers, but also reordering operations, calculating things at compile-time and so on. For a trivial example, see this code: int f() { int sum = 0; for (int i = 0; i < 10; i++) sum += i; return sum; } As you can see from [1], a smart compiler will calculate the sum at compile time and make the function simply return the resulting number (i.e., no loop is generated). If you make "sum" volatile, the compiler is forced to do the loop[2]. [1] https://godbolt.org/z/3sX5mU https://godbolt.org/z/3sX5mU [2] https://godbolt.org/z/F5CiDJ https://godbolt.org/z/F5CiDJ
- aidenn0 7y agoIn terms of "Using volatile too much" I found a comment along the lines of "Not sure why this has to be volatile, but it doesn't work without it" and the answer was "There is a race condition and volatile slows down one path enough to make it go away." Yuck.
- burfog 7y agoThat note at the end about Linux is missing a link to the Documentation/volatile-considered-harmful.txt document. Basically, don't use volatile. Here, with your choice of formatting: https://github.com/torvalds/linux/blob/master/Documentation/process/volatile-considered-harmful.rst https://github.com/torvalds/linux/blob/master/Documentation/... https://www.mjmwired.net/kernel/Documentation/volatile-considered-harmful.txt https://www.mjmwired.net/kernel/Documentation/volatile-consi... https://www.kernel.org/doc/html/latest/process/volatile-considered-harmful.html https://www.kernel.org/doc/html/latest/process/volatile-cons...
- dahfizz 7y agoI think the title and parts of the article are misleading. Using volatile will never make a correct program incorrect. It cannot "break" a correct implementation. It should not be overused, because as the article mentions it makes for slower and more confusing code, but it's not quite something to be afraid of either. It is slower to use volatile, and bad form
- klingonopera 7y ago> "Side note: although at first glance this code looks like it fails to account for the case where TCNT1 overflows from 65535 to 0 during the timing run, it actually works properly for all durations between 0 and 65535 ticks." From example 1, ignoring device and setup-specifics what to do when TCNT1 overflows, it actually works properly for all ticks, both "first" and "second" are unsigned (therefore behaviour is defined), and the delta between them both is always between 0 and 65535, no matter what values they may have, and also correct in all cases. E.g.: timeDelta = timeStampNow - timeStampLast = 0 - 65535 = 1
- klingonopera 7y agoCan someone please tell me why this got downvoted? It's particularly annoying when correcting common misconceptions to get penalised for it... (and if what I said was wrong, I'd also like to know why!)
- icedchai 7y agoProbably because the article already says it works for all ticks...
- klingonopera 7y agoYes, upon rereading I see that, too. When I first read it, I understood in such a way that it only would work until (a timestamp of) 65535 ticks since device-startup had passed, but he was referring to durations of that length. Thank you for clarifying.
- tntn 7y agoBut if the duration is > 65535 ticks, the calculated duration will be wrong, no? There is no mechanism to count how many times TCNT1 overflows, so it will be incorrect if the duration of what you are timing exceeds 65535 ticks.
- klingonopera 7y ago
- SAI_Peregrinus 7y agoThe entire section on declarations can also be fixed by always binding type modifiers and quantifiers to the left. Rewriting the examples: int* p; // pointer to int int volatile* p_to_vol; // pointer to volatile int int* volatile vol_p; // volatile pointer to int int volatile* volatile vol_p_to_vol; // volatile pointer to volatile int This method always starts with the most basic type, then adds modifiers sequentially. The modifier binds to everything left of it.
- kaetemi 7y agoVolatile seems quite sufficient for a PleaseExitThread boolean.
- legohead 7y agoI've never had to use volatile in code. This was all very interesting! For issue #5, a possible solution not mentioned could be to write inline assembly, no? It would keep the array non-volatile and should be portable.
- kelnos 7y agoInline assembly is basically the definition of non-portable.