7 ms·
> the DEC Alpha seems like it should have some sort of cheap fast atomics, but you can't cheaply do Acquire-Release so how about we invent something else ? con
by brickteacup 3y ago
> the DEC Alpha seems like it should have some sort of cheap fast atomics, but you can't cheaply do Acquire-Release so how about we invent something else ?
consume was not invented for Alpha. rather consume should have less overhead than acquire on any platform that needs additional memory barriers to implement acquire semantics (on most platforms consume requires no barriers at all since memory ordering is enforced as a side-effect of dependency ordering. Alpha is an exception to this, i.e. consume does require a barrier on Alpha).
it's also worth noting that rcu_dereference in linux is based on the same idea as consume (using dependencies to enforce memory ordering) the the basic idea is definitely useful. it's more that the C++ standard's specification of consume turned out to be unimplementable in practice
- ajross 3y ago> consume was not invented for Alpha. rather consume should have less overhead than acquire on any platform that [emphasis mine] needs additional memory barriers to implement acquire semantics Um... which platform?[1] Probably the biggest reason that memory ordering semantics are such a mess, and lockless algorithms such an infeasible disaster, is the propensity of people to argue about this stuff in generic terms, using language and concepts from egghead standards writers instead of examples of real machines. Out of order execution is an engineering technique for physical devices. We should talk about the machines first, and the abstractions later. It's like going to the mechanic with a leaky valve and having the explanation come back in terms of cycle efficiencies and isentropic losses. It's not actually helpful to understanding the problem unless you already understand the theory. And you won't ever learn the theory unless you see the problem it's solving. [1] It's a serious question, btw. I genuinely don't know what the target hardware is for "consume" myself, and suspect this is just a pet theory that snuck into the standard.
- reitzensteinm 3y agoAre memory ordering semantics a mess? I implemented simulators for x86, ARM and C++ memory models and the documentation for all three was fantastic. What's part of it is a mess? I'm also not sure I understand your point about physical machines. Surely what we care about are the guarantees each platform gives. There are plenty of cases where the documented worst case is worse than you'll ever be able to replicate - for now. Sticking to the contract means you're safe on next year's machines, too.
- ajross 3y agoTo be glib, the part that answers the question "Who needs consume and why is it there?". As pointed out, Chen answers this incorrectly in the linked article. An upthread poster pointed out (correctly, though they skipped the justification) that Alpha needs this, but was then corrected (incorrectly, AFAICT, though I'm willing to be educated) that this feature was not for Alpha. To be more serious: anyone who's ever tried to implement a lockless algorithm knows how hard this stuff is even if you have a clear and unambiguous set of tools to use to do it. And the C++ memory model is, as mentioned, kinda ambiguous. At the end of the day, in practice these problems are solved on real hardware with real tools that look like "MFENCE" or "DSB" (either directly, or because you're reading the generated assembly trying to figure out how it's going wrong). That's exactly the opposite of the way a good abstraction is supposed to work.
- reitzensteinm 3y agoI'm just not sure what consume and the Alpha have to do with each other. Consume is an attempt to give you a lightweight load when all you want are guarantees related to a dependent load. And the Alpha is the only machine I'm aware of where dependent loads don't give you any guarantees. As far as lockless algorithms go - using something like Loom for Rust greatly simplifies correctness. It has a relatively complete implementation of the C++ memory model, and will catch bizarre edge cases you'd struggle to ever replicate on a real CPU. There are also solvers for C++ that help you do similar reasoning. For _performance_, on the other hand, it's the wild west. I experimented with pairing my home grown system similar to Loom with a MESI cache simulator, testing cache behaviour with lockless algorithms banging on the same cache lines. It showed some promise, but was quite difficult to make ergonomic for the consumer.
- ajross 3y ago> I'm just not sure what consume and the Alpha have to do with each other. This seems like semantic evasion? Consume exists[1] precisely so that you can write code using the C++ memory model that works correctly on a superscalar Alpha chip (not the 21064/21164, those are in-order). Saying that it doesn't have anything to do with alpha seems really strained. If alpha didn't exist then consume wouldn't exist and people would be writing these as unconstrained/normal loads because that works everywhere else. [1] Again, AFAIK. Lots of folks in this thread, including you, seem to be implying that other such hardware exists. But I'm not aware of it. And my broader point is that this kind of "standards first" discussion obscures understanding.
- brickteacup 3y ago> Um... which platform? most platforms with weaker-than-x86 memory models need additional barriers for acquires e.g. on ARM you need LDAR > Probably the biggest reason that memory ordering semantics are such a mess are they though? release/acquire semantics are actually intuitive once you get used to them and are very well suited for pointer-based lockfree data structures. consume is maybe more subtle but it's mostly an optimization on top of acquire, so not very difficult conceptually
- jcranmer 3y agoThe point of memory_order_consume is to be able to omit the load barrier on architectures with weak memory orderings that don't need them for data-dependent operations on the load. This is basically every architecture other than Alpha (which would still need the barrier with memory_order_consume) and x86 and SPARC (which don't need the barrier with memory_order_acquire). ARM and PPC, for example, fall into this category.
- gpderetta 3y ago>Um... which platform? In practice all non-TSO machines that need a load-load barrier for acquire but no barrier for data dependencies (i.e. loads through pointers). Notably this includes ARM and POWER. These architectures still needed relatively expensive fences on the release path. On TSO machines (SPARC, x86 for example) acquire loads and release stores are free. Alpha was the strange outlier that also needed a barrier for the dependent load path. But acquire and release are such a sweet spot for programming that ARM (and apparently POWER) have since added cheap (although not free) acquire load/release stores. Far from being a mess, the C++11 memory model (and acquire/release in particular) has been an unmitigated success that has allowed writing, discussing and proving the correctness of very practical lock free algorithms on real hardware while not resorting to the expensive SEQ-CST model. Also since it inception, all relevant architectures have formalized their memory model and made sure that they can implement acq/rel efficiently.