3 ms·
Exactly. The race condition problem from the example below still exists even without any caching, and perhaps still exists even without any replica. > Imagine
by throwdbaaway 4y ago
Exactly. The race condition problem from the example below still exists even without any caching, and perhaps still exists even without any replica.
> Imagine that after shuffling Alice’s primary message store from region 2 to region 1, two people, Bob and Mary, both sent messages to Alice. When Bob sent a message to Alice, the system queried the TAO replica in a region close to where Bob lives and sent the message to region 1. When Mary sent a message to Alice, it queried the TAO replica in a region close to where Mary lives, hit the inconsistent TAO replica, and sent the message to region 2. Mary and Bob sent their messages to different regions, and neither region/store had a complete copy of Alice’s messages.
A solution to this problem is to update Alice's primary message store to be "region 2;region 1" for a period of time, and get Alice to retrieve messages from both message stores during that period.
However, since you have a caching layer that doesn't have an explicit TTL, such that any cache inconsistency can last indefinitely, you now have a much bigger distributed data problem.
- throwdbaaway 4y ago> 1. The cache tried to fill the metadata with version. > 2. In the first round, the cache first filled the old metadata. > 3. Next, a write transaction updated both the metadata table and the version table atomically. > 4. In the second round, the cache filled the new version data. Another way to look at this is that while the database can handle writes to the 2 tables atomically, the caching layer can't do a consistent read from the 2 tables, because it is either too expensive to do the 2 reads from within a database transaction, or too expensive to do a single read with join? Our current system also suffers heavily from distributed inconsistency, mostly due to these "too expensive" constraints imposed by an ex-FB guy, even though our scale is nowhere near FB level. Yes I am bitter.
- uvdn7 4y agoYou assessment is good. > Our current system also suffers heavily from distributed inconsistency Despite of people having different definitions of cache invalidation (mine is narrower, and a subset), I hope the techniques covered here can be helpful (even a tiny bit).