3 ms·
Directory vs snooping just means you either know exactly which other agents have an interest in a line so you can target them directly, or that you don't so you
by throwawaylinux 4y ago
Directory vs snooping just means you either know exactly which other agents have an interest in a line so you can target them directly, or that you don't so you broadcast and other agents have to "snoop" requests and respond to relevant ones.
The logical coherency protocol that is (somewhat) above that level is what would dictate whether you can forward modified data between CPUs or it would have to write back and be fetched from memory.
AFAIK protocols don't tend to have the writer publish its modification to other owners for two major reasons. First because you then would no longer own it exclusive and would have to go off-chip if it wanted to modify it again immediately afterwards. Second because problems with supporting atomic operations and preventing deadlocks in the protocol can become unmanageable if you could have other owners while you are writing lines. I believe generally the transfer of data is demand-driven, so core 1 wanted to read the line, it would ask core 0 for it. And that modifications will always remove access permission from other cores before the modification completes, so core 0 would have send an intent to write message to core 1 in this case before its store could reach coherency.
Barriers don't necessarily drive any coherency on a cache line level. Caches are always coherent (ignoring access by incoherent agents, where the coherency has to be managed at a software level). Barriers provide ordering between more than one memory access.
The MESIF page is a good basic intro to concepts, naturally state of the art has moved beyond that. Here for example https://course.ece.cmu.edu/~ece742/f12/lib/exe/fetch.php?media=le_power6.pdf https://course.ece.cmu.edu/~ece742/f12/lib/exe/fetch.php?med... a 15 year old CPU has 13 states in its coherency protocol. It's difficult to find a lot of public information about this stuff. Suffice to say these things are so complex that mathematical proofs are used to ensure correctness, no deadlocks, etc. You are right right that data including updates can go directly between agents rather than via memory.
But no matter how smart the hardware gets, the advice to programmers is still the same: stores to a line will slow down access to that line by any other agent, so sharing read-only lines between CPUs is fine, but don't add a lot of writes to the mix.