3 ms·
The tool that monitors cache consistency is easy to build. That by itself doesn't solve anything major. The most important contribution is a novel approach on c
by uvdn7 4y ago
The tool that monitors cache consistency is easy to build. That by itself doesn't solve anything major. The most important contribution is a novel approach on consistency tracing that helps find out "why" caches are inconsistent – pinpoint a bug is much harder than saying "there is a bug". This is based on the key insight that a cache inconsistency can only be introduced (for an invalidation-based cache) in a short time window after the write/mutation, this is what makes the tracing possible.
> How is that solving a computer science problem?
It depends on your definition of a computer science problem. I am definitely not solving P = NP. By your definition, does Google's Paxos Made Live paper solve a computer science problem?
The claim is more a play on Phil Karlton's quote, as the work here makes cache invalidation much easier (in my opinion). Also Phil Karlton's quote doesn't necessarily _make_ a problem a computer science problem, don't you think? I think it's a good quote and there's a lot of truth in it.
- latchkey 4y agoI disagree with your claim that "cache invalidation might no longer be a hard thing in computer science" for the same reason that you say that pinpointing bugs is harder than saying "there is a bug". You can't just build a tool and declare a whole problem space in computer science (and a long running joke) as being solved. > the work here makes cache invalidation much easier The tool itself doesn't make cache invalidation easier, it makes finding bugs (that happen to be cache invalidation bugs), easier. By that logic, if one never wrote a bug, cache invalidation would still be as difficult as before. Again, good work on the tool. That's fantastic. I've definitely done some huge work in this area myself and struggled a lot. Let's also realize that it is also FB internal, so whatever solutions you've come up with aren't really helping anyone else without spending the same amount of engineering time and resources on the problem.
- uvdn7 4y agoYou are right. The assumptions I made are - the definition of cache invalidation https://news.ycombinator.com/item?id=31676102 https://news.ycombinator.com/item?id=31676102 - subsequently, by that definition, _make_ cache consistent in production is the harder problem I understand if you disagree with these premises. And all these make sense. Let's discuss the "cache invalidation problem" by your definition. E.g. in its most generic form, a cache can store arbitrary materialization from any data source. Now when updating the data source, in order to keep caches consistent, you essentially need to transact (cross system transaction) on both the data source and cache(s). Usually cache has more number of replicas, I am not sure running this type of transactions is practical at scale. What happens if we don't transact on both systems (the data source, and cache)? Well, now whenever the asynchronous update pipeline performs the computation, it's done against a moving data source (not a snapshot of when the write was committed). Now let's say the data source is Spanner, which provides point-in-time snapshots. On Spanner commit you can get a commit time (TrueTime) back. Now using that commit time, to read the data and compute cache update asynchronously can be done. Now this does assume whatever we cache (the query e.g.) needs to be schematized, and made known to the invalidation pipeline (in the form of some control plane metadata). I think it's a very fair assumption to make. As otherwise (anyone can cache anything without the invalidation pipeline knowing at all), it's pretty obvious that this problem can't be solved.