6 ms·
It's obviously working well for you but I'm curious how you deal with the circular buffer cache with ATS. For a large library that would seem to be an immediate
by eggnet 11y ago
It's obviously working well for you but I'm curious how you deal with the circular buffer cache with ATS. For a large library that would seem to be an immediate disqualifier. Or maybe I'm not understanding how that works.
I'm assuming as a top 5 site you probably deal with a large library, and appear to be ok with ATS despite that. Comments?
- cbsmith 11y agoCaches generally need to be bounded to be effective. What am I not getting here?
- skuhn 11y agoATS uses a tornado cache, where the write pointer just moves to the next object no matter what. So the disk cache doesn't work as a LRU in the same way as other cache servers. The benefit is that writing is fast and it's constant time since you don't have to do an LRU lookup to pick a place to store the object. The downside is that you are creating cache misses unnecessarily. It's never really been a problem for me in practice. If you have a lot of heartache over it, I would suggest putting a second cache tier in place. Very unlikely to strike out on both tiers.
- cbsmith 11y ago> The downside is that you are creating cache misses unnecessarily. Statistically, it balances out just fine. It turns out that just by controlling how objects get in to the cache, you can effect cache policy enough that eviction policies don't much matter, or at least, a "random out" isn't much different from a "LRU".
- lhedstrom 11y agoTo avoid unnecessary cache writes, there's also a plugin that does implement a rudimentary LRU. Basically, you have to see some amount of traffic before being allowed to get written to the cache. This is typically done in a scenario where it's ok to hit the parent caches, or origins, once or a few times extra. It can also be a very useful way to avoid too heavy disk write load on SSD drives (which can be sensitive to excessive write wear, of course). See https://github.com/apache/trafficserver/tree/master/plugins/experimental/cache_promote https://github.com/apache/trafficserver/tree/master/plugins/...
- cbsmith 11y agoYeah, with SSD's I wonder how much that really helps to improve performance vs. just no cache. Most SSD's have a lot of caching implemented internally, so disk cache can often be self defeating.
- donavanm 11y ago"It Depends." If youre doing "random" writes down to the block dev, like updating a filesystem, it can be very bad. You'll end up hitting the read/update/write cell issues and block other concurrent access. In general I'd worry (expect total throughput to go down, and tail latency way up) around a 10-20% write:read ratio. Conversely if youre doing sane sequential writes, say log structure merges with a 64-256KB chunk size, Id expect much less impact to your read latencies.
- cbsmith 11y agoThis is a read cache though. If you just don't have it, you do reads on the SSD, which are pretty darn quick...
- donavanm 11y agoI believe the term youre looking for is "cache admission policy." This is an adjunct to cache eviction, both are needed for success. I'm very curious what a highly efficient insertion policy and trivial "eviction" policy (FIFO) would look like in practice. Cache insertion research is generally focused on use cases like small associative hardware caches. There's very little applicable public research for larger software caching systems, that Ive found. Probably the best would be Gil Einziger. He appears to have found it as an application of his work on extremely space/time efficient counting of sets, http://www.graduate.technion.ac.il/Theses/Abstracts.asp?Id=28439 http://www.graduate.technion.ac.il/Theses/Abstracts.asp?Id=2... and http://www.cs.technion.ac.il/~gilga/ http://www.cs.technion.ac.il/~gilga/. Of notable mention is TinyLFU http://www.cs.technion.ac.il/~gilga/TinyLFU_PDP2014.pdf http://www.cs.technion.ac.il/~gilga/TinyLFU_PDP2014.pdf. Gil submitted it to Caffeine (Java caching library) last summer, https://github.com/ben-manes/caffeine/pull/24 https://github.com/ben-manes/caffeine/pull/24. It got some traction over the winter and is now showing up in other places (https://issues.apache.org/jira/browse/CASSANDRA-10855 https://issues.apache.org/jira/browse/CASSANDRA-10855). In fact Ben Manes just had a guest post on High Scalability the other day http://highscalability.com/blog/2016/1/25/design-of-a-modern-cache.html http://highscalability.com/blog/2016/1/25/design-of-a-modern.... PS: If anyone is interested in these problems, We're Hiring. edit: https://aws.amazon.com/careers/ https://aws.amazon.com/careers/ or preferably drop me a line to my profile email or my username "at amazon.com" for a totally informal chat (Im an IC, not manager nor recruiter nor sales)
- lumanaughty 11y agoThe Tornado Cache (FIFO) hasn't really been an issue as an eviction algorithm. Most of our caches are sized to hold over a weeks (most over a months) worth of objects in cache. Most objects/traffic is temporal in nature. The popular images and videos are normally only popular for a certain time period. We have looked at not evicting objects on disk if they are in the RAM cache and that has a LRU like eviction algorithm (really it is a CFLUS). Doing this would help in not evicting really popular objects. FIFO has advantages over LRU for disks. It is very efficient with writes since they are all sequential. We use rotational disks when building out very large second tier caches. There are other things to consider when looking at cache in a proxy server. How many bytes does the in memory index take per object in cache (for ATS 10 bytes and that is extremely efficient). Also, does the cache use the filesystem and/or use sendfile for HTTP (like NGiNX), but can't use sendfile when using HTTPS or HTTP/2. Netflix is experience this pain when moving to HTTPS with NGiNX. Every proxy server some advantage, easy of use, well supported APIs, flexible configuration, dynamic loadable modules, HTTP specification compliance, HTTP/2 support, TLS support, performance, etc. It really depends on what you are looking for when choosing a proxy server.
- eggnet 11y agoSounds like if you can size the cache large enough to hold enough data to keep cache miss under control, ATS is solid. Thank you for sharing!
- lhedstrom 11y agoWell, that's true for any cache. The choice of a simple eviction algorithm in ATS is deliberate, and usually yields better cache efficiency than more complex architectures. Fwiw, it does support cache pinning, but that's rarely used nor necessary.
- jjm 11y ago+1 In the long run efficiency is very high. Wonder if a set of test data to demonstrate this as a KPI could be added other than raw benchmark perf.