8 ms·
Transactional Object Storage?
- victorbjorklund 2y agoPretty cool and could be useful for stuff that isnt updated so frequently like a CMS.
- ramesh31 2y agoso... Delta Lake?
- jitl 2y agoThere is also SlateDB, another work in progress take on this. HN link: https://news.ycombinator.com/item?id=41714858 https://news.ycombinator.com/item?id=41714858
- maxmcd 2y agoYeah I think it's very interesting to compare the two. SlateDB expects a single writer and fences writes. This means you can make some serious savings on S3 costs because you're using S3 for consistency but you're batching writes. GlassDB is much more accessible for smaller volume workloads, but gets very costly for high volume because of requests to S3 per-transaction. In-turn the consistency model is easier to reason about because the system is entirely stateless.
- social_quotient 2y agoI found myself thinking about Cloudflare Durable objects new SQLite offering. Nicely detailed here https://simonwillison.net/2024/Oct/13/zero-latency-sqlite-storage-in-every-durable-object/ https://simonwillison.net/2024/Oct/13/zero-latency-sqlite-st... And https://developers.cloudflare.com/durable-objects/best-practices/access-durable-objects-storage/#sql-storage https://developers.cloudflare.com/durable-objects/best-pract...
- mbrt 2y agoThis builds on the same intuition I had, where data can be easily partitioned across objects. What seems to be missing is transactions across different objects though? The flipside is that Cloudflare DO will be a lot faster. Interesting that all these similar solutions are popping out now. I think it would be interesting to combine a SQLite per-object approach with transactions on top of different objects.
- jacobmarble 2y agoIf I had time, I'd like to implement an Iceberg catalog this way.
- akshayshah 2y ago100% this! S3’s lack of write preconditions spawned the whole Iceberg catalog ecosystem anyways.
- jsd1982 2y agoWas it considered to separate each table into its own S3 object?
- jahewson 2y agoOnly if you don’t want transactions across tables?
- mbrt 2y agoIt's a good observation, because I did and decided to keep it out of scope from the base layer. But this is entirely possible. You can wrap GlassDB transactions and encode multiple keys into the same object at a higher level. Transactions across different objects will still preserve the same isolation. The current version is meant to be a base from which to build higer level APIs, somewhat like FoundationDB.
- c4pt0r 2y agoTiDB Serverless is built on S3, it's in production for more than 2 years, Blog link: https://me.0xffff.me/dbaas2.html https://me.0xffff.me/dbaas2.html
- up2isomorphism 2y agoWhenever I saw the claim that “S3 is cheap “, I just cannot take it too seriously.
- mbrt 2y agoYou're right indeed:) but it depends on what you are comparing it with. In this case the comparison is against other managed cloud storage and databases, and in that context I think the claim holds. Is it the cheapest possible storage in existence? No, if you take raw disks and put them in a rack, but I also feel it wouldn't be an entirely fair comparison.
- eek2121 2y agoS3 is one of the most expensive platforms out there, however. Look at backblaze B2 for an example of just HOW expensive S3 is. When i moved from S3 to DO, my bill went from hundreds to $20/mo. The only thing that changed was the hosting provider.
- mbrt 2y agoB2 is mostly S3-compatible, so if they add the same support for preconditions on writes as S3 and GCS, nothing prevents using it as a backend for GlassDB.
- kyle787 2y agoI appreciate the effective use of diagrams. The boundaries are separated really nicely.
- deleted 2y ago[deleted]
- svrakitin 2y agoPretty cool! Do you have any ideas already about how to make it work with S3, considering it doesn't support If- headers?
- choppaface 2y agoNot a full solution, but seeing the OP seeks to be a key-value store (versus full RDBMS? despite the comparisons with Spanner and Postgres?), important to weigh how Rockset (also mainly KV store) dealt with S3-backed caching at scale: * https://rockset.com/blog/separate-compute-storage-rocksdb/ * https://github.com/rockset/rocksdb-cloud Keep in mind Rockset is definitely a bit biased towards vector search use cases.
- mbrt 2y agoNice, thanks for the reference! BTW, the comparison was only to give an idea about isolation levels, it wasn't meant to be a feature-to-feature comparison. Perhaps I didn't make it prominent enough, but at some point I say that many SQL databases have key-value stores at their core, and implement a SQL layer on top (e.g. https://www.cockroachlabs.com/docs/v22.1/architecture/overview https://www.cockroachlabs.com/docs/v22.1/architecture/overvi...). Basically SQL can be a feature added later to a solid KV store as a base.
- tlarkworthy 2y agoYou can do it without using an append only logs https://github.com/endpointservices/mps3 https://github.com/endpointservices/mps3 However it will be much simpler with the new conditional writes
- boulos 2y agoS3 recently added basic matching support (https://aws.amazon.com/about-aws/whats-new/2024/08/amazon-s3-conditional-writes/ https://aws.amazon.com/about-aws/whats-new/2024/08/amazon-s3..., https://docs.aws.amazon.com/AmazonS3/latest/userguide/conditional-reads.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/condit...). They don't have the full suite of GCS's capabilities (https://cloud.google.com/storage/docs/request-preconditions#precondition_criteria https://cloud.google.com/storage/docs/request-preconditions#...) but it's something.
- Onavo 2y agoCongrats on reinventing the data lake? This is actually how most of the newer generations of "cloud native" databases work, where they separate compute and storage. The key is that they have a more sophisticated caching layer so that the latency cost of a query can be amortized across requests.
- mbrt 2y agoIt's my understanding that the newer generation of data lakes still make use of a tiny, strongly consistent metadata database to keep track of what is where. This is orders of magnitudes smaller than what you'd have by putting everything in the same database, but it's still there. This is also the case in newer data streaming platforms (e.g. https://www.warpstream.com/blog/kafka-is-dead-long-live-kafka https://www.warpstream.com/blog/kafka-is-dead-long-live-kafk...). I'm curious to hear if you have examples of any database using only object storage as a backend, because back when I started, I couldn't fin any.
- eatonphil 2y ago> I'm curious to hear if you have examples of any database using only object storage as a backend, because back when I started, I couldn't fin any. Take a look at Delta Lake https://notes.eatonphil.com/2024-09-29-build-a-serverless-acid-database-with-this-one-neat-trick.html https://notes.eatonphil.com/2024-09-29-build-a-serverless-ac...
- mbrt 2y agoWow, not sure how I missed this, but I see many similarities. They were also bitten by lack of conditional writes in S3: > In Databricks service deployments, we use a separate lightweight coordination service to ensure that only one client can add a record with each log ID. The key difference is that Delta Lake implements MVCC and relies on total ordering of transaction IDs. Something I didn't want to do to avoid forced synchronization points (multiple clients need to fight for IDs). This is certainly a trade-off, because in my case you are forced to read the latest version or retry (but then you get strict serializability), while in Delta Lake you can rely on snapshot isolation, which might give you slightly stale, but consistent data and minimize retries on reads. It also seems that you can't get transactions across different tables? Another interesting tradeoff.