7 ms·
> Remember that paths are raw pointers, even in Rust. Most file operations are inherently unsafe and can lead to data races (in a broad sense) if you do not cor
by mjb 4y ago
> Remember that paths are raw pointers, even in Rust. Most file operations are inherently unsafe and can lead to data races (in a broad sense) if you do not correctly synchronize file access.
This is a reality that our operating systems have chosen for us. It's bad. The way Unix does filesystem stuff is both too high level for great performance, and too low level for developer convenience and to make correct programs easy.
What would it look like to go higher level? For example:
- Operating systems could support START TRANSACTION on filesystem operations, allowing multiple operations to be performed atomically and with isolation. No more having to reason carefully about which posix file operations are atomic, no more having to worry about temp file TOCTOU etc.
- fopen(..., 'r') could, by default, operate on a COW snapshot of a file rather than on the raw file on the filesystems. No more having to worry about races with other processes.
- Temp files, like database temp tables, could by default only be visible to the current process (and optionally its children) or even the current thread. No more having to worry about temp file races and related issues.
That sort of thing. Maybe implemented like DBOS: https://dbos-project.github.io/ https://dbos-project.github.io/. Or, you know, just get rid of the whole idea of processes for server-side applications. Was that really a good idea to begin with?
> However, if the locks you use are not re-entrant (and Rust's locks are not), having a single lock is enough to cause a deadlock.
I'm a fan of Java's synchronized, and sometimes wish that Rust had something higher-level object-level synchronization primitive that was safer than messing with raw mutexes (which never seems to end well).
- elij 4y agosame issue with networking (specifically per packet heuristics) and even Java hasn't got that right either.
- surajrmal 4y agoPer application isolated storage is becoming more and more popular. Docker containers, flatpaks, android apps, etc are all providing isolated storage by default. Recursive locks are a nightmare. You cannot reason about lock acquisition order and have to invent new ways to make it safe.
- eska 4y agoThe filesystem is a database. Databases like Postgres implement such a transaction not through a lock, but by keeping multiple versions of the data/file until the transactions using them closed. Yet other databases operate on the stream of changes and the current state of data is merely the application of all changes until a certain time, allowing multiple transactions to use different snapshots without blocking each other (you can parse a file while somebody else edits it and you won’t be interrupted). I read about various filesystems offering some of these features, but not in IO APIs.
- giovannibonetti 4y agoI remember a discussion a while ago to use SQlite as a filesystem engine. I imagine that not needing a server / daemon would make it more reliable for one of the first things needed on boot. However, I don't know what's the recommended way to handle concurrent writes with SQlite. In the end we have a single process handling all the persistence logic, which becomes essentially a server just like Postgres?
- naasking 4y ago> The filesystem is a database A totally shit database designed before we knew anything about databases. Well past time to retire them.
- wongarsu 4y agoThey are ok databases, optimized for very different use cases than normal databases. If you treat files as blobs that can only be read or written atomically, then SQLite will outperform your datasystem. But lots of applications treat files as more than that: multiple processes appending on the same log file, while another program runs the equivalent of `tail -f` on the file to get the newest changes; software changing small parts in the middle of files that don't fit in memory, even mmapping them to treat them like random access memory; using files as a poor man's named pipe; Single files that span multiple terabytes; etc.
- naasking 4y ago
- jcranmer 4y agoMoving to a fully transactional model for filesystems sounds like it will create more pain than help, especially given the historical tendency of filesystem implementations to cheat on semantics every opportunity they get. > - fopen(..., 'r') could, by default, operate on a COW snapshot of a file rather than on the raw file on the filesystems. No more having to worry about races with other processes. You can go a few steps further. If you have some sort of O_ATOMIC mode on creating a file descriptor, that operates on a separate snapshot in both read and write mode (and in the case of open for write, will atomically replace the existing file with the newly written copy on close), that would have more correct semantics than existing POSIX I/O for most applications. You can't make it default at the syscall level, but most language I/O abstractions could probably make it default. Notably, this can be implemented in the OS kernel without actually touching filesystem implementations themselves, just the way the kernel arbitrates between userspace and the filesystem. Another interesting file mode that fixes many use cases is having some sort of "atomic append" for files that handles the use case of simultaneous readers and writers of log files. Set up the API so that you can guarantee that multiple writers each get to add their own atomic log message (without other synchronization), and that readers can never see just part of a log message. The difficult part of the filesystem to reason about is the file metadata and directory listings. I've seen suggestions in the past to use file descriptors instead of strings as the basis for handling directory queries in language implementations. If you don't do at least that much, it's basically impossible to solve TOCTOU issues, but I'm not sure there's a lot you can do with COW snapshotting of directories that still manage to be acceptable in performance. > Temp files, like database temp tables, could by default only be visible to the current process (and optionally its children) or even the current thread. No more having to worry about temp file races and related issues. Creating unnamed temporary files is a solved problem. Named temporary files is more difficult, but create-only-if-doesn't-exist exists on all platforms, and most language APIs have some way of getting to that (it's "x" flag in C11's fopen).
- gpderetta 4y ago>Another interesting file mode that fixes many use cases is having some sort of "atomic append" for files that handles the use case of simultaneous readers and writers of log files My understanding is that O_APPEND gives exactly this guarantee. Readers can read partial records, but with proper record marking it shouldn't be an issue.
- jstimpfle 4y ago> Temp files, like database temp tables, could by default only be visible to the current process (and optionally its children) or even the current thread. No more having to worry about temp file races and related issues. This is one area where the Unix design is much closer to good compared to other historic approaches. The separation of file identity from file paths (not sure if it is a a Unix invention) pretty much allows temp files to be implemented, in fact I believe O_TMPFILE in Linux allows that. > START_TRANSACTION can be implemented on top. Transactions are quite involved, so I wouldn't blame a random file system for not implementing them. Use a database (yes, most of them skip the buffering and synchronisation layer of the filesystem they are stored on, and use e.g. O_DIRECT). > fopen(..., 'r') could, by default, operate on a COW snapshot Something like that would require the definition of safe states in the first place (getting an immutable view is pointless if the current contents aren't valid), so we're almost back at transactions. I think Unix file systems are fine and usable as a common denominator to store blobs, like object stores do, but also easily allow single writers to modify and update them. As ugly and flawed and bugged as most file systems are, most of the value comes from the uniform interface. You need better guarantees than that, you probably have to lose that interface and move to a database.
- deleted 4y ago[deleted]
- wongarsu 4y ago> Operating systems could support START TRANSACTION on filesystem operations, allowing multiple operations to be performed atomically and with isolation Windows supports that since Windows Vista [1], and at least back then it was used by Windows Update and System Restore [2]. But programming languages usually only expose the lowest common denominator of file system operations in easy APIs, so approximately nobody uses it. Also, it's deprecated by now. But maybe someone here has deeper insights in what went right and wrong with that implementation, because on the surface it looks like it would erradicate entire classes of security bugs if used correctly. 1: https://learn.microsoft.com/en-us/windows/win32/fileio/about-transactional-ntfs https://learn.microsoft.com/en-us/windows/win32/fileio/about... 2: https://web.archive.org/web/20080830093028/http://msdn.microsoft.com/en-us/magazine/cc163388.aspx https://web.archive.org/web/20080830093028/http://msdn.micro...
- crabbone 4y agoI don't know the answer, but I'd suspect the following: nobody in storage business cares about what Microsoft is doing beside Microsoft themselves. From storage perspective, supporting Microsoft is always a huge pain. Most of those who do provide support try to limit it to exposing SMB server. I had a misfortune to try to expose iSCSI portal to Windows Server. Luckily, the company was in its relatively early stages where they could decide what things they want to support, and after couple months of struggle they just decided to forget Windows existed. So, I think, this feature, just like WinFS, and probably ReFS after it, will just end up being cancelled / forgotten. The kind of users who use Windows aren't sophisticated enough to want advanced features from their storage, but support and development of this stuff is costly and demanding in other ways. If they cannot sell this as a separate product that runs independently of Windows, there's no real future for it. The only real future for Windows seems to be Hyper-V, but hypervisors, generally, have very different requirements to their storage than desktops. Simpler in some ways, but more demanding in other ways. So, bottom line: it probably didn't see any use because the audience wasn't sophisticated enough to want that, and they couldn't sell it to the audience that was sophisticated enough, but was also smart enough not to buy from Microsoft.
- AceJohnny2 4y ago> Also, it's deprecated by now. Aaaand there it is. Such paradigms need years, at least a decade, to percolate into the general programming culture. Microsoft (and I daresay companies in general) just doesn't have the stability to shepherd such changes.
- pclmulqdq 4y agoThe "file" abstraction sucks. However, it's so deeply ingrained in everything we do that it's nearly impossible to do anything about it. At the language level, if your files don't behave the way Unix files do, then you will have questions from your users, and things like ported databases will not work. Mutexes and locks are also a really tricky API, but making them re-entrant has performance and functionality costs - you basically need to store the TID/PID of the thread that holds the lock inside the lock. I'm sure there's a crate for a mutex-wrapped datatype, trading speed for ease of use, but if not, it's likely a very easy crate to put together.
- dietrichepp 4y agoI agree that it sucks but I do think that it’s possible to solve most of the problems. On Linux, there are some new APIs you see crop up every once in a while. For apps running in the data center, you can use network storage with whatever API you want, instead of just using files. And on the Mac, there have been a couple shifts in filesystem semantics over the years, with some of the functionality provided in higher-level userspace libraries. Solving the problems for databases seems really hard, just based on all the complicated stuff you see inside SQLite or PostgreSQL, and the scary comments you occasionally come across in those codebases. But a lot of what we use files for is just “here is the new data for this file, please replace the file with the new data”, and that problem is a lot more tractable. Another use case is “please append this data to the log file”. Solving the use cases one-by-one with different approaches seems like a good way forward here.
- naasking 4y ago> The way Unix does filesystem stuff is both too high level for great performance, and too low level for developer convenience and to make correct programs easy. Indeed. Imagine if on program start, each process gets its own sqlite database rather than a file system. So many issues would be solved. A user shell is then just another sqlite database that indexes process databases so you can inspect them using a standard interface.
- eqvinox 4y agoLike many other things, the file system is a layer that, after some passing of time, is proving a bit raw in the abstraction it provides. And thus layers are created above it. The same thing is happening pretty much everywhere in computer engineering. Our network protocols provide higher level abstractions (encryption, RPC calls, CRUD, …). Higher level graphics rendering libraries proliferate. Programming languages provide additional layers and safety guarantees. > What would it look like to go higher level? File systems aren't "bad" and don't need to be changed or replaced — we just need to use the higher abstractions that already exist much more. Use a proper database when it's appropriate. Or some structured object storage system, maybe integrated into your programming language. Ultimately, accessing the file system should more and more become akin to opening a raw block device, a TCP socket without an SSL layer, or drawing individual pixels. Which is to say: there are absolutely good reasons to do so, but it shouldn't be your default. And it should raise a flag in reviews to check if it was the appropriate layer to pick to solve the problem at hand. (added to clarify:) This all is just to say: it's more helpful to proliferate existing layers above the file system, than to try to change or extend the semantics of the existing FS layer. Leave it alone and put it in the "low-level tools" drawer, and put other tools on your bench for quick reach.
- mjb 4y ago> This all is just to say: it's more helpful to proliferate existing layers above the file system, than to try to change or extend the semantics of the existing FS layer. Leave it alone and put it in the "low-level tools" drawer, and put other tools on your bench for quick reach. Yes! But it's easier said than done when one of these things is in the stdlib and the other isn't.
- eqvinox 4y ago> one of these things is in the stdlib and the other isn't Oh it's much worse than that. One of these things is what the user sitting in front of their computer has a nice integrated UI to view and search… if a photo editing applications starts storing my photos in a database that I can't easily and simply copy some photo out of, I'll be rather annoyed. And each application having their own UI to do this isn't the solution either, really. [EDIT: there was some stuff here about shell extensions & co. It was completely besides the point. The problem is that the file system has become and unquestionably is the common level of interchange for a lot of things.] …didn't Plan 9 have a very interesting persistence concept that did away with the entire notion of "saving" something — very similar to editing a document in a web app nowadays, except locally? Either way I don't know jack shit about where this is going or should go. I'm a networking person, all I can tell you for sure is to use a good RPC or data distribution library instead of opening a TCP socket ;).
- rowanG077 4y ago> I'm a fan of Java's synchronized, and sometimes wish that Rust had something higher-level object-level synchronization primitive that was safer than messing with raw mutexes (which never seems to end well). I'm not a Java expert. But from looking it up it seems like Java's synchronized is worse in every way to Rusts Mutex(And you most of the time shouldn't even be using mutexes in Rust). You forget a synchronized annotation? Too bad. Whereas a mutex in Rust protects the access to a variable. You literally can't modify it (well modulo unsafe) unless you lock the mutex.
- adql 4y agoPersonally I think keeping it low level but predictable would be far saner approach. App knows what it want to do with data, filesystem can't. Just trimming all of the inconsitencies (especially the "fun" around sync and cached access..) and it's already much better. Of course probably very hard to do without fucking up existing apps ;/ > Operating systems could support START TRANSACTION on filesystem operations, allowing multiple operations to be performed atomically and with isolation. No more having to reason carefully about which posix file operations are atomic, no more having to worry about temp file TOCTOU etc. And your performance goes absolutely to trash once you start trying to make filesystem act as transactional database. Especially if you then try to use the filesystem for a database. > fopen(..., 'r') could, by default, operate on a COW snapshot of a file rather than on the raw file on the filesystems. No more having to worry about races with other processes. Another "performance goes to trash" option. Sure, now more palatable with NVMes getting everywhere, but still a massive footgun > Temp files, like database temp tables, could by default only be visible to the current process (and optionally its children) or even the current thread. No more having to worry about temp file races and related issues. Already an option O_TMPFILE > O_TMPFILE (since Linux 3.11) > Create an unnamed temporary regular file. The pathname argument specifies a directory; an unnamed inode will be created >in that directory's filesystem. Anything written to the resulting > file will be lost when the last file descriptor is closed, unless the file is given a name. ... > Specifying O_EXCL in conjunction with O_TMPFILE prevents a temporary file from being linked into the filesystem in the >above manner. (Note that the meaning of O_EXCL in this case is differ- > ent from the meaning of O_EXCL otherwise.)
- thefaux 4y ago> Already an option O_TMPFILE This is not portable.
- karatinversion 4y agoWell obviously, this is a discussion of the limitations of posix. (Of course, as we see time and again, the real portability solution is for users to standardise on a single platform)
- AtNightWeCode 4y agoFile systems should never be transactional. That is plain incorrect. You or your tech should solve that if needed. The easiest way to see that this is incorrect is that you need to handle the error anyway. Meaning, adding something that can cause another error for an error is not viable. GO got this totally right.
- wtetzner 4y agoThe problem with the application solving the problem is that it can't control other applications attempting to write to the same files.
- AtNightWeCode 4y agoData stores are poor integration points.
- wtetzner 4y agoSure, but the filesystem is already an integration point. I much prefer standard file formats that I can use with any piece of software than trying to get pieces of software to talk in some other way.
- nixomose 4y agorust is supposed to be a systems programming language, so... yeah expect to have systems programming problems to deal with. you want the safety of high level languages, use one...
- grumpyprole 4y ago> I'm a fan of Java's synchronized, and sometimes wish that Rust had something higher-level object-level synchronization primitive There are two big problems with Java's synchronized that I can think of. It adds a monitor and therefore memory overhead to every single object. And it publically exposes the monitor allowing potentially bugs due to unintended use by clients.
- deleted 4y ago[deleted]