3 ms·
> Summary: This is a system which couples your application logic with your persistence layer (in the same OS process). I don't think the paper is talking about
by continuations 7y ago
> Summary: This is a system which couples your application logic with your persistence layer (in the same OS process).
I don't think the paper is talking about the persistence layer. Rather it's focusing on the cache layer.
From the paper:
"Modern internet-scale services often rely on remote, in-memory, key-value (RInK) stores such as Redis and Memcached. These stores serve at least two purposes. First, they may provide a cache over a storage system to enable faster retrieval of persistent state. Second, they may store short-lived data, such as per-session state, that does not warrant persistence."
If you look at Fig.2 in the paper, the database (persistence) layer is always there. What the paper is advocating is to move from an out-of-process cache server like Redis to an in-process cache library like Java Caffeine
- zaroth 7y agoSo for the canonical use case of session state in a load balanced server farm, is the session state being replicated across every single one of these in-process caches? Or, if a request is routed to a server without a copy of the session state, does the in-process cache have to go find a peer server with a copy of the session state? This seems like an easy “solution” to show how it speeds up the happy path while not fully addressing the real-world tradeoffs. Stateless application servers come and go without a care in the world. You can reason very simply about the cost and effects of bringing them up and down. Add in a KV store with sharding, replication, a discovery protocol, a heartbeat protocol, a sync and recovery protocol, etc... and put it all in-process on every application server? Have fun monitoring and debugging this. I much prefer the 3-tier system of a local RAM cache, a Redis cache, and a persistence layer. As much state as possible is idempotent so you can cache it locally as well as in Redis. The load balancer can best-effort route back to the same app server, and your KV lookups automatically check local RAM first. But the central KV store handles replication and is easy to monitor and scale out if needed.
- lallysingh 7y agoHow's in-process issuing the same library much different than using a separate process? What's complicated there in monitoring/debugging that's different that's not otherwise? I think that debugging with one less process is a good thing. And monitoring with one less process boundary to traverse is easier as well.
- _mog1 7y agoThe paper assumes that you have a good amount of related context. I'd recommend at least reading the Slicer paper mentioned several times (https://www.usenix.org/system/files/conference/osdi16/osdi16-adya.pdf https://www.usenix.org/system/files/conference/osdi16/osdi16...).
- inlined 7y agoFor me the most important part was in the abstract: autoscaling. It’s not just about going to zero but going to hundreds of thousands. I can remember my two most painful times with memcache: 1. A client scaled rapidly & we had repeated hits on the same key. We were saturating the NIC in the memcache server. We had to addd indirection for some data and introduced an local in-memory layer in front of memcache. 2. We had to extend the ring size to give memcache better CPU utilization. We had to teach all our code to handle a migration process (reads were fall through to the new ring and writes purged both). We couldn’t just turn off memcache because a total dump would overwhelm the database.