3 ms·
I've been using "sfence" memory barrier to enforce visibility in my programs on amd64/x86_64. It's good to know that my programs might need to be adjusted. Ch
by samsquire 3y ago
I've been using "sfence" memory barrier to enforce visibility in my programs on amd64/x86_64.
It's good to know that my programs might need to be adjusted.
Chris Wellons [1] has an interesting post about using an explicit __ATOMIC_RELEASE and then __ATOMIC_ACQUIRE to tell thread sanitizer that there is definitely a happens-before relationship between threads.
[1]: https://nullprogram.com/blog/2022/10/03/ https://nullprogram.com/blog/2022/10/03/
I have an idea: I would really like to layer an abstraction ontop of the memory model to formulate a protocol for interprocess/interthread communication. In other words, the cache coherency protocol IS the protocol for message passing. No locks!
My understanding is that memory_order_seq_cst or _ATOMIC_SEQ_CST in gcc C means all atomic operations in the same thread happens sequentially from the perspective of other threads.
This whitepaper ("How to miscompile programs with “benign” data races") discusses more about compiler optimisations breaking data races.
[2]: https://www.usenix.org/legacy/event/hotpar11/tech/final_files/Boehm.pdf https://www.usenix.org/legacy/event/hotpar11/tech/final_file...
- gpderetta 3y ago> I would really like to layer an abstraction ontop of the memory model to formulate a protocol for interprocess/interthread communication. In other words, the cache coherency protocol IS the protocol for message passing. No locks! that would be just a queue, right?
- samsquire 3y agoA queue is part of an event loop solution but I'm thinking from the perspective of Erlang, Smalltalk and Dbus and distributed objects communication. Rust async semantics is pretty hard and difficult to define and understand, there is pinning, runtimes, async/await state machines and awaiting and polling and type system interactions. It's all to create a protocol that allows for concurrency and parallelism that is also safe. It seems the cache coherency protocol IS what we're looking for, which is safe happens before relationships and no data races.
- u320 3y agoRust async semantics are complex because they need to support resumable computations on a single C-style stack. It has nothing to do with concurrency, which was already supported, but at the cost of multiple stacks (either green threads at one point or OS threads). Futures just happen to be a very prominent example of resumable computations. Generator functions is another.
- gpderetta 3y agoI'm not sure how cache coherency would help here. Cache coherency is specifically about keeping cache coherent (i.e. not stale) in shared memory systems. It is not applicable to pure message passing systems. Also cache coherency per se it has little to say about sequencing and happens-before. That's the job of the memory model built on top (again, still for shared memory). At an higher level still, the general idea of consistency is of course applicable to both shared memory concurrency and message passing concurrency. Maybe consistency models is what you are thinking about? Alternatively, maybe you want to implement shared-memory on top of message passing. In this case cache coherency is definitely applicable, but I'm not sure how amenable are common cache coherency protocols to be implemented in software (but it is certainly possible, and probably what some RDMA system do).
- samsquire 3y agoI'm not an expert but here's my intuition. The L1 cache is very fast and each core has its own 32 kilobytes or so cache. If we could send between L1 caches of cores then couldn't we have latency below 100 nanoseconds?
- gpderetta 3y agoYou can already have sub 100ns latency on core-to-core communication on some common machines. And yes, thank if the magic of bus snooping the transfer can in practice happen cache to cache, although counterintuitively it is not necessarily what you want for message queues.
- toast0 3y ago
- nick_ 3y ago> I have an idea: I would really like to layer an abstraction ontop of the memory model to formulate a protocol for interprocess/interthread communication. In other words, the cache coherency protocol IS the protocol for message passing. No locks! If I'm understanding you correctly, this is exactly what I've wanted for a while. Let me use the cached values of things when I want, and let me read and write from/to RAM when I want. Those calls imply the load/store order semantics (I think?).
- bcrosby95 3y agoThis sounds like volatile in Java, and while sometimes useful it is still a minor part of writing concurrent code. In particular, once you read the value is potentially stale. If you write to something based upon it you now have a race condition.
- KMag 3y ago> I have an idea: I would really like to layer an abstraction ontop of the memory model to formulate a protocol for interprocess/interthread communication. In other words, the cache coherency protocol IS the protocol for message passing. No locks! That's generally what you get with RCU/lockfree data structures.