4 ms·
High-performance read-through cache for object storage
- toomuchtodo 1y agohttps://old.reddit.com/r/databasedevelopment/comments/1nh1goo/cachey_a_readthrough_cache_for_s3/ https://old.reddit.com/r/databasedevelopment/comments/1nh1go...
- _1tan 1y agoCan someone explain when this would be good solution? We currently store loads of files in S3 and directly ingest them on demand in our Java app API pods. Seems interesting if we could speed up retrievals for sure.
- thinkharderdev 1y agoThe basic tradeoff is that you are paying an extra tax on all requests that are not served by the cache, so you something like this would help if you are reading the same data repeatedly. So, for example, a database built on object storage or something like that.
- Havoc 1y agoIf you can have the cache onsite then it'll likely benefit many things just by virtue of not going through a slow internet link
- onethumb 1y agoThis looks super interesting for single-AZ systems (which are useful, and have their place). But I can't find anything to support the use case for highly available (multi-AZ), scalable, production infrastructure. Specifically, a unified and consistent cache across geos (AZs in the AWS case, since this seems to be targeted at S3). Without it, you're increasing costs somewhere in your organization - cross-AZ networking costs, increased cache sizes in each AZ to be available, increased compute and cache coherency costs across AZs to ensure the caches are always in sync, etc etc. Any insight from the authors on how they handle these issue on their production systems at scale?
- jimbohn 1y agoAssuming the "Designed for caching immutable blobs", I guess the approach is to indeed increase the cache size in each AZ or eat the cross-AZ networking costs.
- shikhar 1y agoYes, that's how we are running it at s2.dev, auto-scaled per-AZ deployments. https://www.reddit.com/r/databasedevelopment/comments/1nh1goo/comment/ne88g0h https://www.reddit.com/r/databasedevelopment/comments/1nh1go...
- trueismywork 1y agoNot the author but. Its a user side read through cache, so no need for pre-emptive cache coherence as such. But there will be a performance penalty for fetching data under write contention irrespective of whether you have single az/multiple AZ. The only way to mitigate the performance penalty here is to have accurate predictive fetching which works for usage patterns.
- immibis 1y agoNot the author, but my suggestion is to use a real infrastructure provider. You will save tons of money.
- denis_dolya 1y ago[flagged]
- immibis 1y agoThanks ChatGPT
- OutOfHere 1y agoFrankly, any web app I develop has configurable in-memory caching built in to it, so I would rather increase its size than add an extrinsic cache. By keeping my cache internal to my application, it's also easier for me to invalidate keys accurately.
- perbu 1y agoIt's about scalability. If you have 100 instances you really want them to share the cache so you increase hitrate and keep egress costs low.
- OutOfHere 1y ago> If you have 100 instances you really want them to share the cache I think that assumes decoupled compute and storage. If instead I couple compute and storage, I can shard the input, and then I won't share the cache across the instances. I don't think there is one approach that wins every time. As for egress fees, that is an orthogonal concern.
- rwmj 1y agoIt'd be cool to put a simple NBD front end on it with an nbdkit plugin. That'd let you trivially turn the immutable objects into Linux devices or use them as backing for qemu virtual disks. (https://libguestfs.org/nbdkit-rust-plugin.3.html https://libguestfs.org/nbdkit-rust-plugin.3.html)
- shikhar 1y agoCheck out ZeroFS (https://www.zerofs.net https://www.zerofs.net), which is using SlateDB (https://slatedb.io/ https://slatedb.io/) ED: Now I catch your drift, it would indeed be cool. ZeroFS requires a commitment to the SlateDB LSM data format.
- mertleee 1y agos2 is one of the coolest technologies that more people need to be talking about - I'm still begging them to move one layer lower -> turning s2 into an incredible middleware for edge IOT deployments! PLEASE if someone from the team sees this - I would pay so much for a ephemeral object store using your same edge protocol (seen in the sensor example from your blog). Cheers!
- shikhar 1y agoHi mertletee, I'd like to understand the request better, mind dropping me an email? It's in my profile
- paulon 1y agolets go!