3 ms·
I think https://news.ycombinator.com/item?id=31672541 https://news.ycombinator.com/item?id=31672541 might address your question.
by uvdn7 4y ago
I think https://news.ycombinator.com/item?id=31672541 https://news.ycombinator.com/item?id=31672541 might address your question.
- moralestapia 4y agoNo, mine was not a question, it was a statement. You misnamed the problem you are trying to solve.
- uvdn7 4y agoOK. I am here to learn. Say we go with your definition of cache invalidation problem. Do you think that's THE hard part of cache invalidation? Say, you are just caching simple k/v data (no joins, nothing). Are you claiming cache invalidation is simple in that case? Also the reason why I didn't mention the cache invalidation dependencies is that _I believe_ it's a solved problem (see the link above, we do operate memcache at scale). I am happy to discuss if and why it wouldn't work for your case.
- moralestapia 4y ago>Are you claiming cache invalidation is simple in that case? No. >I believe it's a solved problem It's not. Suppose you have source-of-truth A (doesn't really matter if it's a key-value store or whatever, it could even be a function for all intents) and a few clients B1, B2, B3, ... that rely on the data from A. You have to keep them in sync. When should B* check if A has changed? Every time they need it? Every minute? Every hour? Every day? This is the cache invalidation problem; which IMO is not even a problem but a tradeoff, but whatever. Epilogue: With all due respect, you have an unbelievable career, IBM, Autodesk, Google, Box and now Meta. None of those companies would give me five minutes of their time because I am self-taught, yet here we are :)
- uvdn7 4y ago> Suppose you have source-of-truth A (doesn't really matter if it's a key value store or whatever, it could be a function for all intents) and a few clients B1, B2, B3, ... that rely on the data from A. You have to keep them in sync. When should B* check if A has changed? Every time they need it? Every minute? Every hour? Every day? That is the cache invalidation problem. A few things to clarify here what you are referring to as clients are cache hosts (as they keep data). You seem to imply that the cache is running on the client? I was referring to cache servers (think memcache, Redis, etc.), for which the membership can be determined. So on update (e.g. when you mutate A, you know all the Bs to invalidate). Now continuing with your example, with cache running on the client. Assuming we are talking about same concept when we say "client", the membership is non deterministic. Clients can come and go (connect and disconnect as they wish). There are some attempts to do invalidation-based cache on clients, but they are hard because of the reason I just mentioned. So usually client cache is TTL'ed. E.g. the very browser you are using to see this comment has a lot of things cached. DNS is not going to send an invalidate event to your browser. It't TTL based. I guess what I am saying is that cache invalidation rarely applies to cache side cache as far as I know. Maybe you have a different example, which we can discuss.
- moralestapia 4y agoCaches, clients, Facebooks, hosts, Metas, Redises(?), ... all those things don't really matter. What matters is B* reads from A, but how often should that be? That's it, literally. That's the whole problem.
- teraflop 4y agoIt doesn't matter whether the cache is co-located with the "client" that ultimately uses the data. Say A is a database, and B1, B2, B3... are memcached servers. The exact same situation applies. > So on update (e.g. when you mutate A, you know all the Bs to invalidate). But "knowing" this is a big part of what people mean when they say cache invalidation is hard! If the value in B is dependent on a complicated function of A's state, then it may be difficult to automatically determine, for any given mutation to A, which parts of B's cache need to be invalidated. > There are some attempts to do invalidation-based cache on clients, but they are hard because of the reason I just mentioned. So usually client cache is TTL'ed. Given this statement, the original title of this submission is even more baffling. If you recognize that data can be cached in clients, and that invalidating those caches is so hard that most systems -- including yours -- just completely abandon the goal of being able to do it correctly/reliably, then how can you claim your system makes it no longer a hard problem?
- unboxingelf 4y agoEpilogue: With all due respect, you have an unbelievable career, IBM, Autodesk, Google, Box and now Meta. None of those companies would give me five minutes of their time because I am self-taught, yet here we are :) Correlation does not imply causation. I’m also self taught and some of those companies have given me a lot more than 5 minutes of their time. Here’s some unsolicited feedback - I interpreted your comment above as implying you know more than $peer, and it comes off a little snarky with the emoji face. That approach doesn’t demonstrate professionalism to me. On the other end, $peer focuses on the technical challenge. Most of the gig boils down to communicating ideas.
- dang 4y agoPlease don't cross into personal attack. It's not obvious what you were intending by that last paragraph, but it doesn't sound nice, and bringing someone's personal details into an argument is already not ok. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- moralestapia 4y agoYou're right, Although I didn't mean it as an attack it def. didn't add to the discussion, so I'll be more careful onwards.
- dang 4y agoAppreciated!