4 ms·
You are right. The assumptions I made are - the definition of cache invalidation https://news.ycombinator.com/item?id=31676102 https://news.ycombinator.com/item
by uvdn7 4y ago
You are right. The assumptions I made are
- the definition of cache invalidation https://news.ycombinator.com/item?id=31676102 https://news.ycombinator.com/item?id=31676102
- subsequently, by that definition, _make_ cache consistent in production is the harder problem
I understand if you disagree with these premises. And all these make sense.
Let's discuss the "cache invalidation problem" by your definition.
E.g. in its most generic form, a cache can store arbitrary materialization from any data source. Now when updating the data source, in order to keep caches consistent, you essentially need to transact (cross system transaction) on both the data source and cache(s). Usually cache has more number of replicas, I am not sure running this type of transactions is practical at scale.
What happens if we don't transact on both systems (the data source, and cache)? Well, now whenever the asynchronous update pipeline performs the computation, it's done against a moving data source (not a snapshot of when the write was committed). Now let's say the data source is Spanner, which provides point-in-time snapshots. On Spanner commit you can get a commit time (TrueTime) back. Now using that commit time, to read the data and compute cache update asynchronously can be done.
Now this does assume whatever we cache (the query e.g.) needs to be schematized, and made known to the invalidation pipeline (in the form of some control plane metadata). I think it's a very fair assumption to make. As otherwise (anyone can cache anything without the invalidation pipeline knowing at all), it's pretty obvious that this problem can't be solved.