7 ms·
Edit: Looks like the slides had an inaccuracy (see replies). Huh, looks like I learned something today :) I think a good way of summarizing volatile is this s
by flafla2 7y ago
Edit: Looks like the slides had an inaccuracy (see replies). Huh, looks like I learned something today :)
I think a good way of summarizing volatile is this slide from my parallel architectures class [1]:
> Class exercise: describe everything that might occur during the
> execution of this statement
> volatile int x = 10
>
> 1. Write to memory
>
> Now describe everything that might occur during the execution of
> this statement
> int x = 10
>
> 1. Virtual address to physical address conversion (TLB lookup)
> 2. TLB miss
> 3. TLB update (might involve OS)
> 4. OS may need to swap in page to get the appropriate page
> table (load from disk to physical address)
> 5. Cache lookup (tag check)
> 6. Determine line not in cache (need to generate BusRdX)
> 7. Arbitrate for bus
> 8. Win bus, place address, command on bus
> 9. All caches perform snoop (e.g., invalidate their local
> copies of the relevant line)
> 10. Another cache or memory decides it must respond (let’s
> assume it’s memory)
> 11. Memory request sent to memory controller
> 12. Memory controller is itself a scheduler
> 13. Memory controller checks active row in DRAM row buffer.
> (May > need to activate new DRAM row. Let’s assume it does.)
> 14. DRAM reads values into row buffer
> 15. Memory arbitrates for data bus
> 16. Memory wins bus
> 17. Memory puts data on bus
> 18. Requesting cache grabs data, updates cache line and tags,
> moves line into exclusive state
> 19. Processor is notified data exists
> 20. Instruction proceeds
> * This list is certainly not complete, it’s just
> what I came up with off the top of my head.
It's also worth mentioning that this assumes a uniprocessor model, so out-of-order execution is still possible which leads to complications in any sort of multithreaded or networked system (See #5, 6, 7, 8 in the OP article).
I think a lot of the confusion stems from the illusion that a uniprocessor + in-order execution model implies to programmers who have never dealt with system-level code. I think in the future, performant software will require a bit more understanding of the underlying hardware on the part of your average software developer -- especially when you care about any sort of parallelism. It doesn't help that almost all common CS curriculum ignores parallelism until the 3rd year or more.
[1] http://www.cs.cmu.edu/~418/lectures/12_snoopimpl.pdf http://www.cs.cmu.edu/~418/lectures/12_snoopimpl.pdf - the last 2 slides
- keldaris 7y agoI don't understand these slides. The volatile keyword does not magically bypass the mechanism by which modern CPUs write to main memory. Am I missing something, or are they somehow meant to be ironic?
- flafla2 7y agoIt (in theory) should bypass any caches in between physical memory and the CPU. Of course this is compiler/arch/OS dependent so YMMV... The slide is admittedly a bit vague, the point is mostly to convey "lots of complicated things that you probably haven't considered are going on in the background to speed up memory accesses in a uniprocessor model." Keep in mind the class is exploring parallel architectures, and that lecture is about snooping-based cache coherence.
- keldaris 7y agoThe volatile keyword will certainly lead to implications for cache coherency, but it cannot bypass the TLB or somehow magically avoid the need to involve the memory controller. Unless I'm grossly misunderstanding something, a majority of the points on the second slide should also be on the first.
- burfog 7y agoYep. The slide is completely wrong. It is showing low-level architecture details that would be 100% identical between the two cases. Volatile changes nothing on that list. Volatile just makes sure the compiler bothers. Otherwise, a pair of writes to the same memory location could be optimized by eliminating the first write. Volatile makes the compiler do that. Of course, the CPU itself may then do this optimization, so volatile is thus not good enough for IO.
- keldaris 7y ago> It is showing low-level architecture details that would be 100% identical between the two cases. To be as charitable as I can possibly be, the only part that could theoretically make sense is that the compiler could emit non-temporal store instructions to bypass the cache. I know compilers currently don't do that for volatile, but I don't know why.