4 ms·
You answered it yourself, I would think it’s because Java doesn’t have that. Using the stronger form emits stronger instructions that have a bigger performance
by slashdev 3y ago
You answered it yourself, I would think it’s because Java doesn’t have that. Using the stronger form emits stronger instructions that have a bigger performance impact. No point paying for that if you don’t need it. If you’re using this kind of thing, it’s specifically because performance really matters and you want to avoid that synchronized compare and exchange that a mutex uses. Otherwise why wouldn’t you just use a mutex or reader writer mutex.
- loeg 3y ago> Using the stronger form emits stronger instructions that have a bigger performance impact. I mean, I don’t think that's true in this case. Both forms generate the same instruction on x86 (lock cmpxchg for both). Godbolt: https://godbolt.org/z/vcMK9j6WP https://godbolt.org/z/vcMK9j6WP (And it appears to be identical on ARM64 as well.) > If you’re using this kind of thing, it’s specifically because performance really matters and you want to avoid that synchronized compare and exchange that a mutex uses. A full fence is not any less expensive than a seq-cst cmpxchg. Do you have any idea why Java is missing acq-rel cmpxchg? I know they are late to the reasonable memory model party but it seems like an odd choice not to just import the C++ memory model wholesale. It's proved to be a decent model in practice!
- slashdev 3y agoI think we’re talking past each other. What I mean is that if you only need a load fence, that’s a lot cheaper (free on x64 if memory serves, just a compiler barrier) than lock cmpxchg. A mutex uses lock cmpxchg, so you’re arguably worse off than just using a mutex (or a reader writer mutex, which has similar performance characteristics for seldom write workloads.) About C++ using the same instruction for both, that’s up to their implementation. There’s no requirement that the more relaxed atomics actually be more relaxed. Only that the stronger atomics don’t permit more relaxed behavior. There may also be other requirements in the spec, that I’m not aware of, that require using the strong compare exchange. It’s been years since I’ve done any substantial lock free programming, I may be remembering incorrectly.
- Jweb_Guru 3y agoA properly efficient sequence lock (no compare exchange or full acquire fence, relaxed loads for reads--which are less expensive on ARM than acquire loads, for good architectural reasons) isn't actually possible to write in C++ and I believe this was in fact a motivation for why Java atomics are the way they are. See https://www.hpl.hp.com/techreports/2012/HPL-2012-68.pdf https://www.hpl.hp.com/techreports/2012/HPL-2012-68.pdf. I suspect this is why Java added load-load.
- gpderetta 3y agoAs per the paper to express it optimally in c++ it requires the compiler to special case +=0. I don't think any compiler does ot yer unfortunately.
- Jweb_Guru 3y agoCompilers are usually very reluctant to add "surprising" optimizations to atomics, even if they're fully justified by the standard, because (1) there are a lot of compiler bugs and/or holes in the standard that are only exposed after combining multiple optimizations involving atomics, (2) most programs explicitly use atomics only rarely, making it not the most productive thing to optimize, and (3) the programs and libraries that do use them are often reliant on them for synchronization, "liveness," or performance behavior that's not mandated by the standard. Common pathological examples include checks that assume the compiler won't skip assigning values if it can prove there's some obscure possible execution where the intermediate value can never be read, assuming that checking the result of a relaxed load in an if statement will guarantee that accesses run in the if will be sequenced after the read, etc.
- bonzini 3y agoIt has always been possible to write it in C++ since C++11, since acquire fences are cheap enough, but not in Java until loadFence() was introduced in JDK 8. The paper you cite proposed compiling an atomic add-zero to MFENCE+MOV, which is 5 times slower than the rest of the seqlock access, just because fences were not in Java at the time and are "a very difficult-to-use and C++11/C11-specific construct". That part of the paper has never made sense to me, and certainly does not matter now that Java has fences. Just wrap the seqlock into a primitive that hides the fence and call it a day.
- mrkeen 3y ago> If you’re using this kind of thing ... why wouldn’t you just use a mutex or reader writer mutex. >> you want the readers to be able to read multiple values atomically. I read this in TFA, and interpreted it to mean you could read multiple values atomically, which would be an answer to your question - mutexes only protect one thing and you can't combine them atomically. But looking at the code, it seems that this does no better? It only allows reading of multiple values insofar as it protects one hard-coded struct which happens to have multiple fields.
- slashdev 3y agoYou can’t do that without a mutex. This provides a way to read a consistent snapshot of multiple atomic values without a mutex. Obviously you could just use a mutex. This may perform better under some workloads, especially if writes are very rare.