6 ms·
Files have awful semantics in just about every way. We should get rid of them ASAP.
by gue5t 9y ago
Files have awful semantics in just about every way. We should get rid of them ASAP.
- icek 9y agoI'd gladly read up on alternative concepts to file systems if you'd be so kind as to supply some searchable terms.
- jamii 9y agohttp://research.cs.wisc.edu/adsl/Publications/ibench-sosp11.pdf http://research.cs.wisc.edu/adsl/Publications/ibench-sosp11.... > We analyze the I/O behavior of iBench, a new collection of productivity and multimedia application workloads. Our analysis reveals a number of differences between iBench and typical file-system workload studies, including the complex organization of modern files, the lack of pure sequential access, the influence of underlying frameworks on I/O patterns, the widespread use of file synchronization and atomic operations, and the prevalence of threads. Our results have strong ramifications for the design of next generation local and cloud-based storage systems. > The iBench tasks also illustrate that file systems are now being treated as repositories of highly-structured “databases” managed by the applications themselves. In some cases, data is stored in a literal database (e.g, iPhoto uses SQLite), but in most cases, data is organized in complex directory hierarchies or within a single file (e.g., a .doc file is basically a mini-FAT file system). One option is that the file system could become more application-aware, tuned to understand important structures and to better allocate and access these structures on disk. For example, a smarter file system could improve its allocation and prefetching of “files” within a .doc file: seemingly non-sequential patterns in a complex file are easily deconstructed into accesses to metadata followed by streaming sequential access to data.
- zokier 9y agoFile systems are just nosql databases, hierarchical key-value blob stores. There are obviously ton of other ways to model databases that could be used. For the other extreme end I think Oracle DB runs quite happily on raw disks, or at least did so at some point. Of course I'm not sure if parent was meaning files as a way to structure/store data (having that hierarchical blobstore) or as a way to access data (something you `open`, `read`, `seek` etc), as they are slightly different things. For a more real world example, take a look how mainframes, especially AS400 (edit: meant System/360 successors), managed data. At least afaik they fundamentally work on a more structured level.
- dsr_ 9y agoOracle DB's preferred method of data storage is for you to hand it disks for Automatic Storage Management, ASM. It then takes care of replication and storage by itself. In practice, this might be a little more performant but incurs significant manageability costs. If you're a committed Oracle shop, it's worthwhile. If you just want one or two database servers and you already have preferred storage methods, use those. (Or, more realistically, use PostgreSQL.)
- dsr_ 9y agoThings to start with: https://en.wikipedia.org/wiki/Soup_(Apple) https://en.wikipedia.org/wiki/Soup_(Apple) https://en.wikipedia.org/wiki/PRC_(Palm_OS) https://en.wikipedia.org/wiki/PRC_(Palm_OS) https://en.wikipedia.org/wiki/Object_storage https://en.wikipedia.org/wiki/Object_storage
- StillBored 9y agoLook at object and capability machines.. Back before OS's became so homogeneous there were a _LOT_ of ideas that didn't map to the modern concept of a file. Some of these machines still exist. For example the AS400/iSeries doesn't really differentiate between ram/storage with its object storage, which means its a perfect fit for a modern non-volatile RAM machine. The original PalmOS, had as similar concept. Of course all the rage the last couple years are key/value stores, which in old terminology one might call KCD (key, count, data) or to rearrange it a bit CKD, aka the technology used for persistent disk storage on IBM mainframes. This is actually one of the things that has gotten a lot easier on the internet the past few years as book scanners have become more common. There now seems to be an effort to preserve old burroghs/whatever manuals online rather than collecting dust in peoples attics.
- protomyth 9y agoThe Newton had a type of object database call soups in place of the file system. It supported queries and frames. The coolest feature was when you removed storage and it still worked with what was still there.
- posterboy 9y agolink to a rant as to answer why files are a bad abstraction? I remember reading here that its not ammendable to metadata, for one thing.
- gue5t 9y agoTry https://danluu.com/file-consistency/ https://danluu.com/file-consistency/ for one point: the filesystem isn't really even viable for data persistence. The filesystem's "you can only persist uninterpreted bytes" policy means software can't maintain any kind of data invariant across runs; everything has to be revalidated if your process ends. ACLs (e.g. unix permissions) are widely regarded as a mistake. File locking is broken: https://gavv.github.io/blog/file-locks/ https://gavv.github.io/blog/file-locks/ File metadata is easy to accidentally mangle (e.g. atime) and hurts performance (even "relatime" is slower than not causing a write for every read). The filesystem is used both for users to organize their data files and for sharing of machine-interpreted data between programs (e.g. shared libraries and system configuration). Humans need human-readable names, and machines get confused by humans renaming things (and should probably be addressing by content, rather cryptographically or in terms of type signatures or specifications). There are no asynchronous syscalls for interacting with the filesystem itself (e.g. `stat()`; for file contents things are onl slightly better). Probably I'm still forgetting a number of problems, but these come to mind offhand.