6 ms·
How would it be different than the filesystem API? There was a great lightning talk a few years ago that I can't seem to find where the author described an API
by dap 10y ago
How would it be different than the filesystem API?
There was a great lightning talk a few years ago that I can't seem to find where the author described an API for storing blobs in a hierarchical namespace. Of course, halfway through, it became clear that it's just the POSIX API: you can "open" handles to objects, "rename" them, remove them, and so on. You'd end a transaction with "fsync()". (Okay, that one's a little more complicated, but I don't think it's as hard as the OP claims, at least for single files. Multiple files are more complicated, but that problem is intrinsically more complicated.)
- derefr 10y agoTwo key differences: 1. "MVCC": the ability to effectively get a handle on a static copy of the entire filesystem, perform mutations to many objects, and then submit the transaction, at which point it will either commit or rollback depending on whether any of those objects have been modified. Windows actually has this: https://en.wikipedia.org/wiki/Transactional_NTFS https://en.wikipedia.org/wiki/Transactional_NTFS allows for exactly the sort of MVCC I'm talking about—and presents a very different API than the POSIX filesystem one. 2. "Object store": as in, you don't interact with the filesystem by getting writable handles to the file's underlying backing store, where processes can effectively treat a file as persistent shared memory. Instead, you ask for a floating unnamed "buffer" object, fill it up, and then submit it to the filesystem to be atomically stored; or, vice-versa, you ask to retrieve a file, and are passed a buffer handle of the representation of the file at the point in time you asked for it (presumably, for optimization's sake, backed by copy-on-write mmap'ed pages from the latest MVCC-linearized copy of the file.) We actually have this today as well, but in an implicit and half-assed way. Files below some size threshold, in modern filesystems, don't use any FS extents for backing, but are instead stored as part of their filesystem directory-entry node. This basically makes the filesystem into an object store for those files—but without actually exposing any guarantee to userland that those files will be read/written as atomic operations. --- To be clearer, there's a problem with this fs-atop-object-store design, as I've stated it so far: An object store doesn't support every use-case a filesystem does, and a filesystem implemented only in terms of an object-store wouldn't be good for some things because of that. It would be terribly expensive to emulate a block device on top of an object store, for example—you'd have to make each emulated block an object, and mutate blocks by transactionally adding an updated block and removing the old block. This means that it wouldn't make sense to keep mountable read-write disk images on such a filesystem; nor would it make sense to keep database backing stores there. But both of those things are effectively things that take no advantage of existing on top of a filesystem in the first place—they do their own index-building, their own sparse-allocation, their own journaling, etc. Making the filesystem an object store just makes it obvious that for those use-cases, the filesystem itself is pure overhead. So, along with the MVCC object-store, the other part of my hypothetical system is a "buffer store": basically like a logical-volume manager, an API that consumes physical block devices and exposes handles to (cheap) logical resizable block devices. The object store could be implemented in terms of one of those logical block devices, but otherwise would completely ignore the existence of it. The filesystem compatibility layer, on the other hand, would allow you to request that a given filesystem object be backed by a persistent buffer (newly-created logical block device) instead of an object; and then whenever you requested that object from the filesystem, you'd get an IO handle to the logical block device, instead of an IO handle to a transactionally-resubmittable copy-on-write mmap of some object.
- Someone 10y agoIt surprised me to read that Microsoft is considering to remove transactional NTFS because "there has been extremely limited developer interest in this API platform since Windows Vista primarily due to its complexity and various nuances which developers need to consider as part of application development" (https://msdn.microsoft.com/en-us/library/windows/desktop/hh802690%28v=vs.85%29.aspx; https://msdn.microsoft.com/en-us/library/windows/desktop/hh8... linked-to from that Wikipedia page)
- Sanddancer 10y agoThe problem is that fsync() runs outside of the control of the program. There's no way for an application to start a transaction, perform steps x, y, and z, and then end a transaction, rolling back to before step x if there are any failures. For example, suppose you're rotating a an audit log file at the same moment your backup program is running. Your backup program reads the directory, and at that same moment, your rotate script had renamed the file, but had created the new file, but data hadn't been written to the file yet. What does the backup program see? does it see the old file you just renamed? Does it see the new zero-length file? Your backup now has an indeterminable state, and potentially lost data, because the backup received a consistent view of the overall data. Were there a way of creating a transaction, the second program looking at the same data would either see the old file, or would see the new file and the rotated file. This is where a transactional file system would be of great benefit, because it limits the amount of indeterminate state to a very minimum, even while multiple programs are operating on the same file.
- frutiger 10y agoRealistically, backups need to be based on atomic snapshots, like `zfs`.
- Sanddancer 10y agoYep. I'll admit my example was a bit convoluted, but I was trying to show a way in which common tasks can race and create undesirable situations. ZFS snapshots would definitely make things quite a bit more predictable.