4 ms·
`lock ; xadd` isn't really fundamentally different than `ll ; add ; sc; b again` The latter is a bit clunky but the core more or less implements them in the sa
by throwawaylinux 3y ago
`lock ; xadd` isn't really fundamentally different than `ll ; add ; sc; b again`
The latter is a bit clunky but the core more or less implements them in the same way. Acquire a line exclusive, load value, increment it, write it back. And you can hold the line exclusive such that the conditional store failure cause is mostly a formality, and can't actually become an infinite loop.
No general purpose atomics are done by shipping the operation to the cache or to memory controllers, it just doesn't work[*]. So even if they look slightly different in the core, they all end up looking exactly the same at the caches and coherency protocols, and that is where atomics are slow. Well any sharing of cache lines updates really.
[*] EDIT: That is to say it doesn't work for performance, for many reasons. Some CPUs do have "remote atomics" something like that which does exactly this, but they are not intended to be broadly used.
- gpderetta 3y agoYes, AFAIK in practice many architectures special case some ll/sc sequences to guarantee forward progress and fairness.