4 ms·
It's amazing what you can do with S3. It's one of the best things that AWS has to offer. I wonder, is there a formal definition for a set of primitives that al
by techn00 2y ago
It's amazing what you can do with S3. It's one of the best things that AWS has to offer.
I wonder, is there a formal definition for a set of primitives that allow you to build an ACID database? Assume an API of some kind (in this case, S3) that you can interact with - and provides, I don't know, locks, % durability, etc.
What would make you say, 'Having those primitives, I CAN build an ACID database on top of it'?
- rad_gruchalski 2y agoWake me up when s3 supports wrie at offset. Until then it’s all gimmicky. Writing small objects and retrieving them later is very inefficient and costly for large data volumes. One can do roll ups, sure, but with roll ups there’s no longer a way to search through the single rolled up file. One needs some compute to download the complete file and process it outside of s3.
- ethegwo 2y agoRandom reads and sequential writes are enough to build a log-structured database, but S3 does not really support the latter.
- rad_gruchalski 2y agoOne can do sequential writes by simply writing to a chunk objects at a key with the offset in the name. For example, sizes in MBs: tmp/uploads/object-0, tmp/uploads/object-1024, tmp/uploads/object-2048 would be a rolled up object of size 2048MB + whatever is in the object-2048 file.
- MadsRC 2y agoIt’s actually surprisingly efficient if you batch writes at the expense of some added latency. The WarSyream team found that batching into chunks of either 4MB of data or 250ms was optimal. Downside is the 250ms latency. But then again, a fair amount of workloads can deal with 250ms of latency.
- rad_gruchalski 2y agoKafka does batch reading anyway. Have a look at the reader, it’s a loop with read and read timeout. Usually a 100ms per single loop iteration.
- ncruces 2y agoS3 can at least do a multi-part upload, where any given part is a copy of a range of an existent object. Then you can finish the upload overwriting the previous object. GCS, unfortunately, does not support copying a range. OTOH, it has long supported object append through composition. The challenge with both offerings is that writes to a single object, and writes clustered around a prefix, are seriously rate limited, and consistency properties mostly apply to single objects.
- rad_gruchalski 2y agoYeah, but you cannot multipart single chunk into a larger complete file. You need all chunks one way or another. Multipart upload starts and ends from all chunks. GCS and Azure support this too. S3 does a maximum of 1k objects, GCS 32 objects, and Azure blob storage, afair 5k objects. Both can do an operation similar to what you described for S3 with various alternatives of read at offset + length and rolling those up. In all cases, you end up always rolling up into a new key that isn’t available for read until roll up is done. It’s kinda useless for heavy write scenario. Compare that to your normal fs operation. Write at offset to an existing file with size smaller that offset will just truncate the file to the offset, and continue writing.
- ncruces 2y agoYou can? You create a multi-part with 3 chunks: the 1st part is a copy range of the prefix, the 2nd part the bit you want to change, and the 3rd a copy range of the suffix? And yes, all of this is useless for heavy (and esp. concurrent) writes.
- rad_gruchalski 2y agoWe both said the same thing. You kinda can but cannot. Yes, you can replace some part of an existing object but you cannot resize it, not can you do anything parallel with that. So you kina can but cannot. And this trick will work in gcs and azure, here you have to move the new object to an old key yourself after the roll up. But why not while you’re already at it.
- deleted 2y ago[deleted]
- n00j 2y agoI would guess something like Apache Iceberg would be something close to this? https://iceberg.apache.org/ https://iceberg.apache.org/ Is a table format which can be use via trino, spark, flink, java apis, pything API?
- ethegwo 2y agoWe are building tonbo: https://github.com/tonbo-io/tonbo https://github.com/tonbo-io/tonbo , An embedded KV database allows to use S3 as storage backend, and we are trying to implement SQLite virtual table on it: https://github.com/tonbo-io/sqlite-tonbo https://github.com/tonbo-io/sqlite-tonbo a real pay-as-you-go DB.
- grapesodaaaaa 2y agoThat’s really cool! I’m personally really interested in serverless DB offerings. I’m not sure if yours scales well, but I always seem to hit the limits of a single RDBMS instance at some point as a product matures. There are plenty of ways to scale out traditional RDBMS, but serverless offerings make it so easy to scale out.
- ethegwo 2y agoThanks, easy-to-scale is the first thing we consider, also using S3 as a shared storage service makes architecture easy to achieve this.
- arianvanp 2y agoThis blog post is great in explaining this: https://blog.mbrt.dev/posts/transactional-object-storage/ https://blog.mbrt.dev/posts/transactional-object-storage/
- justincormack 2y agoYou basically just need a log [1]. Sirupsen goes through the basics in [2] (video) I did a talk about building on object storage with lots of examples [3] (video) [4] (slides with links) [1] https://engineering.linkedin.com/distributed-systems/log-what-every-software-engineer-should-know-about-real-time-datas-unifying https://engineering.linkedin.com/distributed-systems/log-wha... [2] https://www.youtube.com/watch?v=RFmajOeUKnE https://www.youtube.com/watch?v=RFmajOeUKnE [3] https://www.youtube.com/watch?v=ei0wwTy6_G4 https://www.youtube.com/watch?v=ei0wwTy6_G4 [4] https://static.sched.com/hosted_files/kccncna2024/c8/Object%20%20Storage%20is%20all%20you%20Need.pdf https://static.sched.com/hosted_files/kccncna2024/c8/Object%...
- ComputerGuru 2y agoHaving a consistent log is sufficient with an atomic compare operation is sufficient for a distributed database but its performance will be extremely questionable. CAS is always the slow step, in this case, pathologically slow. The magic is to do whatever you can to avoid it until absolutely necessary. The availability of consistent, ordered, synchronized timestamps across all nodes is something most distributed databases require as a prerequisite. How you handle violations of that (and to what degree of accuracy you can rely on it) make a considerable difference. Depending on how you structure the underlying pages, you’ll get to decide how availability at the log level translates to availability in your user/app-facing interface and whether you will end up sacrificing consistency, availability, or partition tolerance. Basically, S3 with its recent consistency guarantees and all-new CAS support is sufffiicent in-and-of itself. But for anything other than the most basic (least amount of data, lowest frequency writes, etc) you’ll need a considerable amount of magic to make it useable. The most straightforward approach would be to use the existing whole of another database but swap out the backend and then tweak the frontend accordingly. SQLite lets you use custom vfs providers (already used to provide fairly efficient SQLite over http without serving the entirety of the database, but previously not for writes) and with Postgres you can use foreign data wrappers. But in both cases you’ll basically have to take out a lock to support writes, either on a page or a row (either risk lots of contention or introduce a ton of locking and network latency overhead).