5 ms·
The article is well written, but I am annoyed at the attempt to gatekeep the definition of a filesystem. Like literally any abstraction out there, filesystems
by YouWhy 3y ago
The article is well written, but I am annoyed at the attempt to gatekeep the definition of a filesystem.
Like literally any abstraction out there, filesystems are associated with a multitude of possible approaches with conceptually different semantics. It's a bit sophistic to say that Postgres cannot be run on S3 because S3 is not a filesystem; a better choice would have been to explore the underlying assumptions; (I suspect latency would kill the hypothetical use case of Postgres over S3 even if S3 had incorporated the necessary API semantics - could somebody more knowledgeable chime in?).
A more interesting venue to pursue would be - what other additions could be made to the S3 API to make it more usable on its own right - for example, why doesn't S3 offer more than one filename per blob? (e.g., a similar to what links do in POSIX)
- bilalq 3y agoThis might be of interest to you: https://neon.tech/blog/bring-your-own-s3-to-neon https://neon.tech/blog/bring-your-own-s3-to-neon. There's also the OG Aurora whitepaper: https://www.amazon.science/publications/amazon-aurora-design-considerations-for-high-throughput-cloud-native-relational-databases https://www.amazon.science/publications/amazon-aurora-design...
- zX41ZdbW 3y agoClickHouse can work with S3 as a main storage. This is possible because a table is a set of immutable data parts. Data parts can be written once and deleted, possibly as a result of a background merge operation. S3 API is almost enough, except for cases of concurrent database updates. In this case, it is not possible to rely on S3 only because it does not support an atomic "write if not exists" operation. That's why external, strongly consistent metadata storage is needed, which is handled by ClickHouse Keeper.
- afiori 3y agoIs a "write if not exists" atomic operation enouhg as a concurrency primitive for database locks?
- justincormack 3y agoYes, its not necessarily the most efficient mechanism (could be a lot of retries) but its sufficient. See the Delta Lake paper for example [0] [0] https://people.eecs.berkeley.edu/~matei/papers/2020/vldb_delta_lake.pdf https://people.eecs.berkeley.edu/~matei/papers/2020/vldb_del...
- yencabulator 3y agoWhen talking about analytical databases for "big data", yeah. They generally just want a "atomically replace the list of Parquet files that make up this table", with one writer succeeding at a time. That would not be a great base to build a transactional database on.
- mlhpdx 3y agoConditional PUT would be a great addition to S3, indeed.
- buremba 3y agoThat would probably require them to rewrite a non-trivial part of S3 from scratch.
- yencabulator 3y agoGoogle Cloud Storage supports create-if-not-exist and compare-and-swap on generation counter. S3 is much harder to use as a building block without tying your code into a second system like DynamoDB etc. https://pkg.go.dev/cloud.google.com/go/storage#Conditions https://pkg.go.dev/cloud.google.com/go/storage#Conditions
- 8n4vidtmkvmk 3y ago[flagged]
- deleted 3y ago[deleted]
- jillesvangurp 3y agoThe notion of postgres not being able to run on s3 has more to do with the characteristics of how it works than with it not being a filesystem. After all, people have developed fuse drivers for s3 so they can actually pretend it's a filesystem. But using that to store a database is going to end in tears for the same reasons that using e.g. NFS for this is also likely to end in tears. You might get it to work but it won't be fast or even reliable. And since NFS actually stands for networked file system, it's hard to argue that NFS isn't a filesystem. Whether something is or isn't a filesystem requires defining what that actually is. A system that stores files would be a simple explanation. Which is clearly something S3 is capable of. This probably upsets the definition gatekeepers for whatever more specific definitions they are guarding. But it has a nice simple logic to it. It's worth considering that file systems have had a long history, weren't always the way they are now, and predate the invention of relational databases (like postgres). Technically before hard disks were invented in the fifties, we had no file systems. Just tapes and punch cards. A tape would consist a single blob of bits, which you'd load in memory. Or it would have multiple such blobs at known offsets. I had cassettes full of games for my commodore 64. But no disk drive. These blobs were called files but there was no file system. Sometime, after the invention of disks file systems were invented in the early sixties. Hierarchical databases were common before relational databases and filesystems with directories are basically a hierarchical database. S3 lacking hierarchy as a simpler key value store clearly isn't a hierarchical database. But of course it's easy to mimic one simply by using / characters in the keys. Which is how the fuse driver probably fakes directories. And S3 even has APIs to listfiles with a common prefix. A bigger deal is the inability to modify files. You can only replace them with other files (delete and add). That kind of is a show stopper for a database. Replacing the entire database on every write isn't very practical.
- buremba 3y agoNeon.tech runs Postgresql runs on S3. They persist the WAL to S3 so that they can replicate the data and bring it to local ssds I assume.
- mike_hearn 3y agoWell, RocksDB never overwrites files except the manifest which is small. And you can write DB features on top of that. So that's an example of a database that can work with the S3 limitations.
- defaultcompany 3y agoI’ve wondered this also because it can be handy to have multiple ways of accessing the same file. For example to obfuscate database uuids if they are used in the key. In theory you could implement soft links in AWS by just storing a file with the path to the linked file. But it would be a lot of manual work.