6 ms·
Fuchsia has an interesting take on filesystems [1]. One can write it completely in the user-space, avoiding expensive kernel<-->user-space switching. Additional
by zenlibs 7y ago
Fuchsia has an interesting take on filesystems [1]. One can write it completely in the user-space, avoiding expensive kernel<-->user-space switching. Additional benefit of storage sand-boxing comes for free, as each app can implement it's own fs, with the rest of the system unawares of it's existence.
I wish such a fully-user-space option existed for Linux. This work is philosophically in the opposite direction, moving more functionality into kernel space for perf benefits.
[1]: https://fuchsia.dev/fuchsia-src/the-book/filesystems.md https://fuchsia.dev/fuchsia-src/the-book/filesystems.md
- blattimwind 7y agoFS hooking (ala usvfs) is very similar to what you want.
- comex 7y agoThat document describes filesystems which are accessed over IPC, not filesystem-as-library like you seem to be describing. In fact, it's the same basic idea as FUSE. One user process (accessing the filesystem) makes an IPC call to another user process (server that implements the filesystem), which necessarily passes through the kernel and performs a context switch in each direction. On the other hand, it's quite possible that Fuschia's IPC is better optimized than FUSE, so it might have better performance in practice.
- zenlibs 7y agolibfs [1] is a userspace library offered by fuchsia abstracting the traditional vfs (virtual filesystem interface), allowing the fs to exist wholly in userspace, without a kernel component. Quoting: > Unlike more common monolithic kernels, Fuchsia’s filesystems live entirely within userspace. They are not linked nor loaded with the kernel; they are simply userspace processes which implement servers that can appear as filesystems [1]: https://fuchsia.googlesource.com/fuchsia/+/master/zircon/system/ulib/fs/ https://fuchsia.googlesource.com/fuchsia/+/master/zircon/sys...
- wahern 7y agoYou said, "avoiding expensive kernel<-->user-space switching", which is wrong. Filesystems are implemented entirely in user space, just not in the same user space processes. Consumers exist in separate processes from the producers--plural, because the underlying block device storage may be managed by processes separate from the processes managing VFS state. Context switches are a necessary part of having separate user space processes, and context switching through kernel space (or at least some protected, privileged context) is necessary in order to authenticate messaging capabilities. Note that there are ways to minimize the amount of time spent in privileged contexts. Shared memory can be used to pass data directly, but unless you want all your CPUs pegged at 100% utilization the kernel must be involved somehow to optimize IPC polling. In any event, the same strategies can be used for in-kernel VFS services, so it's not a useful distinction.
- zenlibs 7y agoThe only reason to use an IPC or go through the kernel is to expose the fs to the rest of the OS. If an app doesn't intend to expose the fs, the entirety of the fs can exist within the app process. Quoting: > "Unlike many other operating systems, the notion of “mounted filesystems” does not live in a globally accessible table. Instead, the question “what mountpoints exist?” can only be answered on a filesystem-specific basis -- an arbitrary filesystem may not have access to the information about what mountpoints exist elsewhere."
- wahern 7y agolibfs is an abstraction layer around some of the VFS bits. An analogous Unix approach be would be shifting the burden of compact file descriptor allocation (where Unix open(2) must return the lowest numbered free descriptor) to the process rather than the kernel. (IIRC this is also actually done in Fuschia as part of its POSIX personality library.) Notice in the above that it's implied that that actual filesystem server (e.g. that manages ext4 state on a block device) is in another process altogether. And so for every meaningful open, read, write, and close there's some sort of IPC involved. A process accessing a block device directly without any IPC is something that can already be done in Unix. For example, you can open a block device in read-write mode directly and mmap it into your VM space. Also, see BSD funopen(3) and GNU fopencookie(3)[1], which is a realization of similar consumer-side state management, except for ISO C stdio FILE handles; it's simpler because ISO C doesn't provide an interface for hierarchical FS management. There's no denying that Fuschia's approach is more flexible and the C interface architecture more easily adaptable to monolithic, single-process solutions. But it stems from it's microkernel approach which has the side effect of forcing VFS implementation code to be implemented as a reusable library. There's no reason a Unix process couldn't reuse the Linux kernel's VFS and ext4 code directly except that it was written for a monolithic code base that assumes a privileged context. Contrast NetBSD's rump kernel architecture where you can more easily repurpose kernel code for user space solutions; in terms of traditional C library composition, NetBSD's subsystem implementations have always fallen somewhere between Linux and traditional microkernels and so were naturally more adaptable to the rump kernel architecture. [1] See also the more limited fmemopen and open_memstream interfaces adopted by POSIX.
- Palomides 7y agoon linux, something like intel's SPDK will let you do everything, very quickly, in userspace
- monocasa 7y agoWhich doesn't quite have the same semantics, as you can't multiplex the disk easily. Your one app totally owns the device.
- monocasa 7y agoExcept you ultimately want your buffer cache, virtual memory, and vfs tightly coupled because they're three sides of the same coin. IMO, the ideal combo looks something like this FUSE/BPF work combined with XOK's capability based buffer cache rather than trying to split everything out into user mode.
- wahern 7y agoThis all gets back to the microkernel debates. As you say, it's far easier to implement such an architecture by stuffing a lot of the most important bits into a monolithic kernel. But various microkernel projects have shown this isn't necessary, and doing things this way has proven brittle and insecure. So-called safe languages don't help, either, because the whole purpose of doing this in a monolithic, shared memory context is precisely because it's easier to move fast and break things in terms of unsafe optimizations (e.g. circular, direct pointer references) unburdened by careful, formal constraints. If written in Rust every other line of Linux code would be wrapped in unsafe{}. Most of the parts that needn't be would be better moved into user space, anyhow.
- monocasa 7y agoThere is no uKernel out there that has anything on Linux wrt to FS perf. The uKernels have not shown that they've solved the FS problem as well as monolithic kernels.
- tathougies 7y agoFuchsia doesn't have an interesting take. It's a take copied from microkernels, which I contend are the most written kinds of kernel, not because they're always technically superior, but because OS authors like writing them.
- ori_b 7y ago> avoiding expensive kernel<-->user-space switching And replacing it with expensive userspace<--->userspace switching. But because the kernel is still doing the context switch between the processes, now you're doing userspace<-->kernel<--->userspace switching.
- nybble41 7y agoIn the ideal case the IPC mechanism is just a block of memory shared between processes running simultaneously on separate cores, so there is no need to context-switch. You still want some sort of kernel-based blocking semaphore to avoid busy-waiting when the IPC queue is empty, but since it's only used when the queue is empty it isn't part of the performance-critical path.
- nwmcsween 7y agowelcome to exokernels circa 1994