4 ms·
It might not happen only once in a while. Consider the following pattern in which a Fetch() does a Get and then an Insert on cach-miss: Fetch("1"); Fetch("1
by eis 3y ago
It might not happen only once in a while. Consider the following pattern in which a Fetch() does a Get and then an Insert on cach-miss:
Fetch("1"); Fetch("1"); Fetch("2"); Fetch("2"); Fetch("3"); Fetch("3"); ...
This will result in your queue being 100% filled with items that are marked as hit because of the second Fetch for each key. After the queue is full each first Fetch will result in an eviction (new key to insert) which has to walk the whole queue because the only item marked not-hit is at the head of the queue due to being lazy promoted.
My O(n*m) reply was because your O(m) was considering a whole run over multiple requests and my O(n) was only considering one request. So if I were to also talk about multiple requests then your eviction algorithm needs to look at O(n*m) items for evictions in the example pattern I presented. O(n) per request for m requests.
A malicious actor could abuse this fact to cause a DoS.
No you can't just bound the number of objects it checks because that would result in a flurry of cache misses as it would mean you have to fail the insert.
- 1a1a11a 3y agoConsider the the request sequence you gave a, a, b, b, c, c, d, d ... and a cache size of 2, when the first c arrives, the cache is in the following state (left is head and right is the tail, inserted at the head): b(1) a(1) and it will trigger an eviction, and requires resetting 2 (or O(m)) objects with the following state changes b(1) a(1) -> b(1) a(0) -> b(0) a(0) -> c(0) b(0) the next c will change the bit to 1, and we will have c(1) b(0), When d arrives, b will be evicted, and we have d(0) c(1) Note that we only check/reset one object this time instead of O(m). Therefore, in the worst case between each O(m) eviction, you need O(m) requests to set the m objects to "visited", strictly no more than a LRU cache.
- eis 3y agoOk I see where the difference is now. You are clearing the hit bit for all items marked as hit on a single eviction. I mean yes I guess that means that during the following request you don't have to check all items in the queue but now you have a different problem: you evict items prematurely. With a cache of size 2 I can't show it but consider a cache of size 3 and the following pattern: a, a, b, b, c, b, d, e When d arrives state changes c(0) b(1) a(1) -> d(0) c(0) b(0). When e arrives state changes d(1) c(0) b(0) -> e(0) d(0) c(0). Note how b got evicted from the cache even though it was requested more recently and more frequently than c because d cleared out the hit bit for both b and c. Now you traded off hitrate vs latency. But the worst case latency is still O(n) just can't happen multiple times in a row anymore. Imagine you have a cache of 1 million items all marked as hit and you get a new request which would evict all 1 million items. You gained latency for some subsequent requests but increased it in the current one. Crucially you also just cleared out the hit marker of whole cache.
- 1a1a11a 3y agoI agree with you. Any eviction algorithm would have an adversarial workload, so a simpler algorithm is better because you know what the workload was and whether it happens in production; on the other end, the complicated one only makes this worse and the failure of a cache is often catastrophic. :)
- eis 3y agoI can't agree with that. More sophisticated algorithms don't necessarily have to make things worse. In fact I think the opposite is true. Simple algorithms are the ones most likely to fail catastrophically because they are made for a certain range of access patterns. You don't always know the workload beforehand and can pick the appropriate cache policy. Workloads can even change during runtime. I haven't seen a simple cache algorithm that performs well in every workload and I believe any that wants to succeed needs to incorporate different sub-algorithms for different workloads and be adaptive.
- 1a1a11a 3y agoHave you experienced workload changes that caused an algorithm to be less effective? I am happy to learn about it. From my experience and study of traces collected in the past, I haven't seen such cases. Of course, there is diurnal changes, and sometimes short abrupt changes (e.g., failover, updating deployment), but I have not seen a case that a workload change that needs a different algorithm, e.g., switch from LRU-friendly to LFU-friendly. Having an ideal adaptive algorithm is certainly better, and I do observe different workloads favor different algorithms, but all adaptive algorithms have parameters that are not optimal for a portion of workloads and also hard to tune (including ARC). If you give me an algorithm, I can always find a production workload for you that the algorithm does not work well. "Simple algorithms are the ones most likely to fail catastrophically because they are made for a certain range of access patterns" I do not see how FIFO and LRU fail catastrophically except on scanning/streaming workloads. Happy to learn about it.
- eis 3y ago