8 ms·
Writing a file system from scratch in Rust
- azhenley 6y agoThere’s also this file system chapter from a series on writing an OS in Rust: http://osblog.stephenmarz.com/ch10.html http://osblog.stephenmarz.com/ch10.html
- est31 6y agoAnd for code there is TFS https://github.com/redox-os/tfs https://github.com/redox-os/tfs
- phjesusthatguy3 6y agoWe've attempted this as well and it's not as simple as it seems. The issues we've run into have made us reconsider porting our FS handlers to Rust, although we are cautiously optimistic about later results.
- ianlevesque 6y agoAny more specifics?
- unethical_ban 6y agoI have read down to the implementation section, but for my money, this is the best way to describe the high level function and behavior of a filesystem that I have ever seen.
- ridiculous_fish 6y agoA very accessible (though dated) intro to filesystems is Practical File System Design, by Dominic Giampaolo. PDF link: http://www.nobius.org/practical-file-system-design.pdf http://www.nobius.org/practical-file-system-design.pdf
- Q6T46nT668w6i3m 6y agoFrankly, not too much has changed since Giampaolo. In fact, it is still standard reading in many graduate seminars on the subject!
- vondur 6y agoIs he the guy who did the BeOS filesystem?
- peterkelly 6y agoYes. Later went to Apple and did Spotlight. https://en.wikipedia.org/wiki/Dominic_Giampaolo https://en.wikipedia.org/wiki/Dominic_Giampaolo
- saagarjha 6y agoAnd APFS: https://developer.apple.com/videos/play/wwdc2016/701/ https://developer.apple.com/videos/play/wwdc2016/701/
- deleted 6y ago[deleted]
- dm319 6y agoIt would be nice if the intro had a brief explanation of why a disk needs to be divided into blocks. Otherwise, I really enjoyed this read from the perspective of a lay person.
- unethical_ban 6y agoThe disk / inodes need to know where to start looking for a file's contents, like the address in memory for RAM. Or like the mail: We subdivide by city, then ZIP, then street, then address. So the inode says "The data for my file starts at block 72 and is 3 blocks long" (or something like that). The disk then goes there, and reads blocks 72,73,74. Each block is 4KiB large often, so if you have a 10KiB file, you still take up ceiling(file size/block size) blocks. That's why there is a difference between "File size" and "Size on disk" when you look at disk usage summaries.
- saurik 6y agoThis doesn't explain why blocks are valuable as you could use byte addressing. The reason why blocks are valuable is for similar reasons to why memory is divided into pages. (Which is all I am going to write, as I don't have the time to answer this well today. But "you need to know where things are on disk" isn't an answer.)
- sedatk 6y agoThis isn’t the reason about block addressing at all since it existed before paging was a thing. Addressing storage content in blocks is closely aligned with how direct access storage is organized in “sectors”, usually 512-byte blocks. Disks don’t have byte access interfaces, so it doesn’t make sense to use one. It also lets you to address much larger data structures with limited sized variables. With block addressing, you can access terabytes of storage with a 32-bit integer.
- saurik 6y ago(I am awake now, so I can provide a more useful answer to this question.) > This isn’t the reason about block addressing at all since it existed before paging was a thing. I didn't say "blocks exist because pages exist", I said "blocks exist for similar reasons to why pages exist". How you thought that required temporal causality is beyond me. :/ > Addressing storage content in blocks is closely aligned with how direct access storage is organized in “sectors”, usually 512-byte blocks. OK, great: now you have to answer why, as you are just punting the answer. The answer, FWIW, is similar to the answer for why memory is divided into pages. > Disks don’t have byte access interfaces, so it doesn’t make sense to use one. But, of course, they could have a byte-access interface; in fact, that was in some sense the entire question: why don't they? Note, BTW, that RAM also doesn't have a byte-access interface: you have to access it by page. (And hell: many modern disks are just giant flash memories!) The super high-level API for working with memory provided by the CPU as part of instructions lets you access it by the byte, but again: that's also true of disks and blocks, where the super high-level API for working with files absolutely lets you work with it accessing specific bytes. > It also lets you to address much larger data structures with limited sized variables. With block addressing, you can access terabytes of storage with a 32-bit integer. This is true, and is a partial answer to the original question. Essentially, a filesystem is serving the same problem space as a page table, but instead of mapping virtual pages of processes to physical pages it is mapping block-sized regions of files to blocks on disk. In both cases, it is thereby beneficial to save some bits in that map. However, as pages of memory are usually 4k and blocks on disk are usually 512 bytes--though you see larger for both of these (64k pages are becoming popular)--that really isn't that many bits you are saving. The real reason is that there is a cost to random access: the more entries in your page-table / filesystem block list the more space it takes and the harder it is to find the one you want; and, while you can "compress" ranges, that also comes with its own cost on random access. On disks made out of spinning rust, this random access cost is particularly ridiculous, as it might require rotating a heavy piece of metal and physically adjusting magnets to achieve the specific offset you are looking for. You thereby never want to load a single byte at a time: you want to pull a full block. There is also another interesting cost to random access on writes, which applies to some kinds of memory and applies to most kinds of disks, which is that you just don't have the precision required to write a single byte at a time: think about trying to isolate one byte using a magnet! So the way these things work is that they have to read large subsets of data at one time and then write it back. If you can align your filesystem to match the blocks of a file along these write boundaries, you thereby decrease the amount of work the disk has to do. (And, again: this is very similar to the work that memory managers are doing with respect to attempting to align pages of large allocations to make the virtual memory areas efficient in the page table, which, again, is pretty much a filesystem for memory with processes as files.) You thereby also often find yourself attempting to tune your block/page size to that of the underlying hardware, and if you build higher-level data structures (such as databases using B-trees or dictionaries using RB-trees) you will want to align those at the higher-level as well. (And it is thereby then also fascinating that it took as long as it did for people to realize that if you want to implement an in-memory tree that you should always prefer a B-tree--a data structure designed around disk blocks--over an RB-tree, for all of the same reasons.) Blocks (and pages) thereby exist to help all layers of the system most efficiently take advantage of linear access patterns to get efficient parallel loading and caching to solve a problem that is otherwise subject to high random access costs and potentially-ridiculous fragmentation issues.
- Immortal333 6y agoShameless plug. I did similar in my OS course. But, in C. Github: https://github.com/immortal3/EbFS https://github.com/immortal3/EbFS Warning: Terribly written. many hacks.
- RealityVoid 6y agoSoooo... how does it work? I'm not asking about the structure or how it's organized. I mean... is the filesystem in a file or... how? Background: I mostly do embedded stuff so at a glance I would have expected low level primitives (like, HW interactions, registers and stuff) but I see none. So maybe, my expectation, when tacking a problem, of interacting with the HW directly, does not stand in modern environments. Even better, but unrelated question... how the heck does a x86 OS request data from the HDD?
- mcpherrinm 6y agoYou'd presumably have some "block device" abstraction between your filesystem and your device driver. Don't want to re-implement a FS for each type of hardware. On a Linux system, you can read, eg, /dev/sda1 from userspace, which is what it looks like this filesystem probably does. As for how you actually request data from the hard drive: There's older ATA interfaces, and BIOS routines from them, which I suspect is what most hobbyist OSes would use. A more modern interface is AHCI. The OSDev wiki has an overview, where you can see how the registers work: https://wiki.osdev.org/AHCI https://wiki.osdev.org/AHCI
- pkaye 6y agoLooks like a filesystem in a file.
- keithnz 6y agoas an aside, for our embedded system we use https://github.com/ARMmbed/littlefs https://github.com/ARMmbed/littlefs for our flash file system, it has a bit of a description on its design and its copy on write system so that it can handle random power loss. Be nice to see some of these kinds of libraries done in Nim or Rust.
- 6y ago
- aquabeagle 6y agofrankly, It never helped on my resume but I enjoyed writing. Frankly, if I saw this code on a resume, I'd keep looking. scanf("%s",buffer); // Debug :: printf("hash : %ld\n",hash("cdRoot")); switch(hash(buffer)) { //command : "ls" case 5863588:
- Ericson2314 6y agoMy dream is to add enough type parameters so in-memory collections can also work as (not horribly tuned!) on-disk datastructures. It's a nice ambitious goal which can really drive language and library design.
- bluejekyll 6y agoAlways fun to see this type of work. I notice the usage of OsString, and it made me wonder: does the way an OS encodes it’s strings potentially make this FS non-portable between OSes? If I want to mount a drive formatted with this FS, would the OsString be potentially non-portable? There was a lot of discussion in the past around TFS https://github.com/redox-os/tfs https://github.com/redox-os/tfs, my understanding is that effort has kinda lost steam.
- fiddlerwoaroof 6y agoThis is really cool, I wish someone would fund it.
- still_grokking 6y agoThat dead[1] project? Actually everything around "Redox" looks like: https://gitlab.redox-os.org/redox-os/tfs/issues/66 https://gitlab.redox-os.org/redox-os/tfs/issues/66 [1] https://gitlab.redox-os.org/redox-os/tfs/issues/80 https://gitlab.redox-os.org/redox-os/tfs/issues/80
- qchris 6y agoRedox is still very much active... [1] https://gitlab.redox-os.org/groups/redox-os/-/activity https://gitlab.redox-os.org/groups/redox-os/-/activity [2] https://www.redox-os.org/news/ https://www.redox-os.org/news/
- vinc 6y agoI recently wrote a very simple and naive filesystem in rust for a toy OS I'm building and it was quite an interesting thing to do: https://github.com/vinc/moros/blob/master/doc/filesystem.md https://github.com/vinc/moros/blob/master/doc/filesystem.md Then I implemented a little FUSE driver in Python to read the disk image from the host system and it was wonderful to mount it the first time and see the files! https://github.com/vinc/moros-fuse https://github.com/vinc/moros-fuse
- ravenstine 6y agoIs there any advantage in writing a custom file system for a niche purpose? It seems like most file systems are just different variations of managing where/when files are written simultaneously. Could a file system written specifically for something like PostgreSQL cut out the middle-man and increase performance?
- tene 6y agoYou may be interested in a paper written by the Ceph team: "File Systems Unfit as Distributed Storage Backends: Lessons from 10 Years of Ceph Evolution" https://www.pdl.cmu.edu/PDL-FTP/Storage/ceph-exp-sosp19.pdf https://www.pdl.cmu.edu/PDL-FTP/Storage/ceph-exp-sosp19.pdf There are definitely some significant benefits you can get from managing your own storage, rather than using a filesystem.
- deleted 6y ago[deleted]
- topspin 6y agoYes. Oracle has done this (ASM) to eliminate overhead, implement fault tolerance and provide a storage management interface based on SQL, for example. I once made a 'file system' to mount cpio archives (read-only) in an embedded system. Cpio is an extremely simple format to generate and edit (in code) and mounting it directly was very effective.
- formerly_proven 6y agoI suspect operating on block storage directly may both be easier and more reliable for databases, since about 75 % of the complication in writing transactional I/O software is working around the kernel's behavior.
- zzz61831 6y agoKernel's fsyncing behavior is one thing, but just relying on a massive amount of fragile C code running in kernel is a significant liability, especially if your software is a centralized database and crashes, panics will bring down everything.
- blackrock 6y agoOnce you have the file system, and a scheduler, don’t you have a basic rudimentary operating system? How soon until someone builds an Operating System developed in Rust? Maybe make it microkernel-based this time.
- smt88 6y ago> How soon until someone builds an Operating System developed in Rust? Redox[1] has been around for almost as long as Rust has. I first heard about it 4-5 years ago. They had an interesting competition a while back challenging people to figure out how to crash it. 1. https://www.redox-os.org/ https://www.redox-os.org/
- sjwright 6y agoI'd be curious to experiment with a file system where all of the file and path metadata is centrally stored in a sqlite blob. Is sqlite fast enough for dealing with file system metadata requests?
- shmerl 6y agoSomething like bcachefs could have been written in Rust.