12 ms·
Replacing a cache service with a database
- hoppp 1y agoThe cache service is a database of sorts that usually stores key value pairs. The difference is in persistence and scaling and read/write permissions
- Supermancho 1y agoie A cache is a database. The difference is features and usage.
- hinkley 1y agoA database is usually a union of all of the questions that can be asked about a topic. A cache by definition is a subset of that. Subsets are not the sets. And if you treat them as if they are, which 90% of people do, you’re gonna have a bad time.
- Supermancho 1y ago> A database is usually a union of all of the questions that can be asked about a topic That's some AI level sophism. A database is a durable store of data that can be modified and read. Ostensibly, we're talking about computer databases. You can define the soft terms at your leisure and to suit your needs. There are many categories of discussion that will never intersect with this definition. Communication is not a database. Art is not a database. History is not a database. Medicine is not a database. et al. A cache is a database. Differentiating a cache and database by label is a misnomer.
- hinkley 1y ago> That's some AI level sophism. Oh fuck off. Calling everything AI is so 2024. A database is a system of record. It can also be a source of truth. A cache is neither. Treating it as one is dangerous. Insisting others should is idiocy.
- Supermancho 1y ago> A database is a system of record. It can also be a source of truth. This is meaningless. A cache is used in lieu of the value because it's considered equivalent. > Insisting others should is idiocy. I did no such thing. Good luck with whatever.
- barrkel 1y agoNo, what makes a cache a cache is invalidation. A cache is stale data. It's a latent out of date calculation. It's misinformation that risks surviving until it lies to the user.
- jedberg 1y agoThis is true but a lot of the trouble in invalidation can be avoided by using smarter cache keys. For example, on reddit, fully rendered comments are cached, so that the renderer doesn't have to redo its work. But the cache key includes the date of the last edit on the comment, which is already known when requesting the value from the cache. In this way, you never have to invalidate that key, because editing the comment makes a new key. The old one will just get ejected eventually.
- cbsmith 1y agoSo close to getting push driven architecture...
- phoronixrly 1y agoRails also has a take on this https://github.com/rails/solid_cache https://github.com/rails/solid_cache
- xixixao 1y agoThis is a good deep dive into the complexity around caching: https://stack.convex.dev/caching-in https://stack.convex.dev/caching-in Having caching by default (like in Convex) is a really neat simplification to app development.
- simonw 1y agoA friend of mine once argued that adding a cache to a system is almost always an indication that you have an architectural problem further down the stack, and you should try to address that instead. The more software development experience I gain the more I agree with him on that!
- jitl 1y agoYeah my architecture problem is that Postgres RDS EBS storage is slow as dog. Sure our data won’t go poof if we lose an instance but it’s so slow. (It’s not really my architecture problem. My architecture problem is that we store pages as grains of sand in a db instead of in a bucket, and that we allow user defined schemas)
- jmull 1y agoThat's true in my experience. Caches have perfectly valid uses, but they are so often used in fundamentally poor ways, especially with databases.
- AtheistOfFail 1y agoI disagree. For large search pages where you're building payloads from multiple records that don't change often, it could be beneficial to use a cache. Your cache ends up helping the most common results to be fetched less often and return data faster.
- DrBazza 1y agoI'd argue the database falls into that category. The two questions no one seems to ask are 'do I even need a database?', and 'where do I need my database?' There are alternate data storage 'patterns' that aren't databases. Though ultimately some sort of (Structure) query language gets invented to query them.
- barrkel 1y agoCaches suck because invalidation needs to be sprinkled all over the place in what is often an abstraction-violating way. Then there's memoization, often a hack for an algorithm problem. I once "solved" a huge performance problem with a couple of caches. The stain of it lies on my conscience. It was actually admitting defeat in reorganizing the logic to eliminate the need for the cache. I know that the invalidation logic will have caused bugs for years. I'm sure an engineer will curse my name for as long as that code lives.
- tengbretson 1y agoMaybe these distinctions are useful to people in some situations, but to me this reads like wondering whether we can replace houses with buildings.
- jayd16 1y agoMore like they're stocking the fridge and wondering what living next to the market is like.
- zeras 1y agoI think a fundamental mistake I see many developers make is they use caching trying to solve problems rather than improve efficiency. It's the equivalent of adding more RAM to fix poor memory management or adding more CPUs/servers to compensate for resource heavy and slow requests and complex queries. If your application requires caching to function effectively then you have a core issue that needs to be resolved, and if you don't address that issue then caching will become the problem eventually as your application grows more complex and active.
- chamomeal 1y agoIdk I think caching is a crucial part of many well-designed systems. There’s a lot of very cache-able data out there. If invalidating events are well defined or the data is fine being stale (week/month level dashboards, for example), that’s a fantastic reason to use a cache. I’d much rather just stuff those values in a cache than figure out any other more complicated solution. I also just think it’s a necessary evil of big systems. Sometimes you need derived data. You can even think about databases as a kind of cache: the “real” data is the stream of every event that ever updated data in the database! (Yes this stretching the meaning of cache lol) However I agree that caching is often an easy bandaid for a bad architecture. This talk on Apache Samza completely changed how I think about caching and derived data in general: https://youtu.be/fU9hR3kiOK0?si=t9IhfPtCsSyszscf https://youtu.be/fU9hR3kiOK0?si=t9IhfPtCsSyszscf And this interview has some interesting insights on the problems that caching faces at super large scale systems (twitter specifically): https://softwareengineeringdaily.com/2023/01/12/caching-at-twitter-with-yao-yue/ https://softwareengineeringdaily.com/2023/01/12/caching-at-t...
- hinkley 1y agoThere are a lot of things necessary to be a successful human but doing them without doing the fundamentals just makes you a monkey in a suit. Caching belongs at the end of a long development arc. And it will be the end whether you want it too or not. Adding caching is the beginning of the end of large architectural improvements, because caches jam up the analysis and testing infrastructure. Everything about improving or adding features to the code slows down, eventually to a crawl.
- jayd16 1y agoSo I guess this guy wants Firestore (or the OSS equivalent)?
- eatonphil 1y agoMany of these points are not compelling to me when 1) you can filter both rows and columns (in postgres logical replication anyway [0]) and 2) SQL views. [0] https://www.postgresql.org/docs/current/logical-replication-row-filter.html https://www.postgresql.org/docs/current/logical-replication-...
- avinassh 1y agoIs it possible to create a filter that can work over a complex join operation? That's what IVM systems like Noria can do. With application + cache, the application stores the final result in the cache. So, with these new IVM systems, you get that precomputed data directly from the database. Views in Postgres are not materialized right? so every small delta would require refresh of entire view.
- jamesblonde 1y agoSome of these questions are informed by the Redis/DynamoDB or Postgres/MySQL world the author seems to inhabit. Why would you want to do this? "I don’t know of any database built to handle hundreds of thousands of read replicas constantly pulling data." If you want an open-source database with Redis latencies to handle millions of concurrent reads, you can use RonDB (disclaimer, I work on it). "Since I’m only interested in a subset of the data, setting up a full read replica feels like overkill. It would be great to have a read replica with just partial data. It would be great to have a read replica with just partial data." This is very unclear. Redis returns complete rows because it does not support pushdown projections or ordered indexes. RonDB supports these and distion aware partition-pruned index scans (start the transaction on the node/partition that contains the rows that are found with the index). Reference: https://www.rondb.com/post/the-process-to-reach-100m-key-lookups-per-second-with-rest-api-and-python-clients https://www.rondb.com/post/the-process-to-reach-100m-key-loo...
- miggy 1y agoWe had a critical service that often got overwhelmed, not by one client app but by different apps over time. One week it was app A, the next week app B, each with its own buggy code suddenly spamming the service. The quick fix suggested was caching, since a lot of requests were for the same query. But after debating, we went with rate limiting instead. Our reasoning: caching would just hide the bad behavior and keep the broken clients alive, only for them to cause failures in other downstream systems later. By rate limiting, we stopped abusive patterns across all apps and forced bugs to surface. In fact, we discovered multiple issues in different apps this way. Takeaway: caching is good, but it is not a replacement for fixing buggy code or misuse. Sometimes the better fix is to protect the service and let the bugs show up where they belong.
- andersmurphy 1y agoI guess CPUs are pretty buggy with all their caches. If only the hardware people could fix their buggy systems. In all seriousness sometimes a cache is what you need. Inline caching is a classic example.
- WillDaSilva 1y agoThere are times when a cache is appropriate, but I often find that it's more appropriate for the cache to be on the side of whoever is making all the requests. This isn't applicable when that is e.g. millions of different clients all making their own requests, but rather when we're talking about one internal service putting heavy load on another one. The team with the demanding service can add a cache that's appropriate for their needs, and will be motivated to do so in order to avoid hitting the rate limit (or reduce costs, which should be attributed to them).
- spyspy 1y agoYou cannot trust your clients. Period. It doesn’t matter if they’re internal or external. If you design (and test!) with this assumption in mind, you’ll never have a bad day. I’ve really never understood why teams and companies have taken this defensive stance that their service is being “abused” despite having nothing even resembling an SLA. It seemed pretty inexcusable to not have a horizontally scaling service back in 2010 when I first started interning at tech companies, and I’m really confused why this is still an issue today.
- gethly 1y agoEvent-sourcing is a powerful tool that helps with exactly this. Why spin up a cache server when you can spin up another read DB instance for the same price and get unlimited capabilities...
- valentinammm 1y ago[dead]
- mannyv 1y agoInstead of redis etc you could get away with static files served via a cdn. Again, you should test. But the main reason imo for redis is connections and speed, not just speed.
- chamomeal 1y agoHey OP you may have seen this already, but in case you didn’t see my other comment, you should definitely check out this talk by Martin Kleppman. https://youtu.be/fU9hR3kiOK0?si=t9IhfPtCsSyszscf https://youtu.be/fU9hR3kiOK0?si=t9IhfPtCsSyszscf It details Apache samza, which I didn’t totally grasp but it seems similar to what you’re talking about here. He talks about how if you could essentially use an event stream as your source of truth instead of a database, and you had a sufficiently powerful stream processor, you could define views on that data by consuming stream events. The end result is kind of like an auto-updating cache with no invalidation issues or race conditions. Need a new view on the data? Just define it and run the entire event stream through it. Once the stream is processed, that source of data is perpetually accurate and up-to-date. I’m not a database guy and most of this stuff is over my head, but I loved this talk and I think you should check it out! It’s the first thing I thought of when I read your post.
- shivasaxena 1y ago[dead]
- ajcp 1y agoThank you for sharing. I thoroughly enjoyed the talk and am as well not a "database guy".
- stevoski 1y agoSomething missing from the article: For the type of cache usage described in the article, cache lookups are almost always O(1). This is because a cache value is retrieved for a specific key. Whereas db queries are often more complicated and therefore take longer. Yes, plenty of db queries are fetching a row by a key, and therefore fast. But many queries use a join and a somewhat complicated WHERE clause.
- interstice 1y agoI've been thinking a lot recently about edge/client layer data sync, interesting to hear where others are at. Noria seems to have got as far as a smart way to store and manage tabular data, however this doesn't seem to help much when the frontend is built on blobs & if one isn't prepared to write the additional layer for read/write on top of the rest of the fetching system. The dumb/MVP approach I'd like to try sometime is close-to-client read only sqlite db's that get managed in the background and neatly handled by wrapper functions around things like fetch. The part I've been slowly thinking about is Noria style efficient handling of data structures while allowing for 'raw' queries, ideally I'd like to set this up so the frontend doesn't need an additional layers worth of read/write functionality just to have CDN-like behaviour. Maybe something like plugins to [de/re]normalise different kinds of blob to tables (from gql, groqd, etc). I'd also like to include a realtime cache invalidation/update system to keep all clients in sync without cache clearing... If I ever get that far.
- interstice 1y agoThis got me thinking a bit more. Rest / GraphQL / Groq handled with adapters, flatten anything nested that references an ID to the row level. Opinionated queries (queries only fetch a superset/subset of the same structure). Fetched data 'fans out' the new content into the rows based on ID to fill out/update structure. Lives in a service worker or side by side with frontend. Drops oldest/least fetched data when limits are reached. Would something like that work? Alternatively just ship an entire shallow copy of least changed / most used data as sqlite db's to the edge, push updates to those, and fetch from source anything that isn't in the DB. Might be simpler.