4 ms·
> but atomic reads can technically be optimised out (though it might not currently be done by most compilers). The memory model should explicitly forbid this o
by exDM69 5y ago
> but atomic reads can technically be optimised out (though it might not currently be done by most compilers).
The memory model should explicitly forbid this optimization, especially if you use an explicit memory ordering with your loads and stores. I'm not 100% sure what the default (sequential consistency) does in this case.
And if you do `read(); barrier(); read(); barrier();` like you should on most CPU architectures, the compiler definitely is not allowed to reorder the reads (barriers affect the compiler AND the CPU). Which barrier you need depends on the MMIO semantics and CPU architecture you're on. Adding volatile does not change this but may inhibit other, useful optimization.
In your example of `volatile_read(); volatile_read()`, the compiler is not allowed to optimize or reorder the loads but the CPU is still allowed to reorder them with each other or other instructions. Adding `volatile` does not make this example work.
- masklinn 5y ago> The memory model should explicitly forbid this optimization, especially if you use an explicit memory ordering with your loads and stores. It has no reason to. Reading the same location twice, even atomically, has no “reason” to change the value so reusing the same read is perfectly valid. If there is an other operation between the two and that has a sequencing implication then yes, otherwise no. > In your example of `volatile_read(); volatile_read()`, the compiler is not allowed to optimize or reorder the loads but the CPU is still allowed to reorder them with each other If there are memory reads then sure but i’m assuming actual mmio here (since volatile is otherwise useless) and i’d think an mmio implementation which reorders mmio reads is completely broken.
- exDM69 5y ago> It has no reason to. Reading the same location twice, even atomically, has no “reason” to change the value so reusing the same read is perfectly valid. No, this is definitely not how atomics with memory orderings work. If you do two atomic reads, the CPU will do two load instructions. Neither the compiler or the CPU is allowed to optimize this. > and i’d think an mmio implementation which reorders mmio reads is completely broken. This is exactly how ARM SoCs work. The CPU core does not know if a load address is mmio or not and may reorder the loads as it sees fit. It's up to the memory subsystem to deal with this, the CPU core does not care. I've done lots of mmio code on ARM at $work in the past few years, and we've spent a lot of time to make sure we have all our memory barriers right. They are quite tricky on ARMv8, as mmio can be either on the system bus or the main memory bus which need distinct kinds of memory barriers. Using volatile in mmio code on modern CPUs is almost certainly a bug. Some microcontrollers may be an exception.
- masklinn 5y ago> No, this is definitely not how atomics with memory orderings work. Of course it is. > If you do two atomic reads, the CPU will do two load instructions. Neither the compiler or the CPU is allowed to optimize this. If you perform two atomic loads on the same location in sequence, then it is completely feasible and normal that the two loads would return exactly the same value, thus the second load can be optimised away. Even under sequential consistency, this is perfectly, well, consistent. It's also perfectly valid to fold atomic writes. See http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/n4455.html http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/n445... for more discussion on the subject No compiler currently bothers doing that (that I know), but it's not because they can't, it's because the effort is probably not worth it, at least at the moment. > I've done lots of mmio code on ARM at $work in the past few years, and we've spent a lot of time to make sure we have all our memory barriers right. They are quite tricky on ARMv8, as mmio can be either on the system bus or the main memory bus which need distinct kinds of memory barriers. I can believe that. > Using volatile in mmio code on modern CPUs is almost certainly a bug. Some microcontrollers may be an exception. Not using volatile in mmio code on modern compilers is almost certainly a bug as well.
- comex 5y ago> No, this is definitely not how atomics with memory orderings work. If you do two atomic reads, the CPU will do two load instructions. Neither the compiler or the CPU is allowed to optimize this. Yes they are. And other optimizations are allowed too. If you do an 32-bit atomic load and then mask the result with `& 0xff`, the compiler can theoretically narrow it into an 8-bit atomic load, as long as the architecture guarantees that such a load still provides the necessary ordering guarantees – which I think is true on most architectures. This won't necessarily work for MMIO, where incorrectly-sized accesses may trap or return the wrong data. But the compiler is not responsible for preserving MMIO semantics if you only asked for atomic. That said, I'm not aware of any major compilers performing /any/ significant optimizations on atomics, so you won't get in trouble for this in practice. > This is exactly how ARM SoCs work. The CPU core does not know if a load address is mmio or not and may reorder the loads as it sees fit. It's up to the memory subsystem to deal with this, the CPU core does not care. You can configure in memory attributes whether to allow reordering or not. > They are quite tricky on ARMv8, as mmio can be either on the system bus or the main memory bus which need distinct kinds of memory barriers. Which SoC?