5 ms·
Author here - would love to hear your comments, as I think this is an interesting area of debate, especially with digital obsolescence waiting to bite us all.
by whyleyc 12y ago
Author here - would love to hear your comments, as I think this is an interesting area of debate, especially with digital obsolescence waiting to bite us all.
- sliken 12y agoWell a file is just a special case of an object, one with a single parent (a directory), and a name typically designed to be human readable. One that's particularly limiting metadata. Typically owner, name, and a few time stamps. So sure people want to download streaming videos. But chances are the result will be put into something that lets them view them by director, title, genre, newest, etc. Generally it seems more natural and user friendly to have your photos, documents, music, and videos indexed. That way you don't have to remember the name, directory, or even what you computer you used to access it. You just look it up by when you used it last (i.e. I edited it yesteday) or by whatever metadata you remember. So if a user's main method of access is by index of metadata not available in the filesytem, why even have a file system? A database seems more natural. After all why is /directory/fi lename more important than being able to look based on arbitrary metadata?
- altcognito 12y ago> why even have a file system? None of this implies that the underlying file systems aren't really useful in this situation. In fact, proper use of a file system would include separation of the media indexes from the media themselves, so that if you want to migrate to a better index, you can do so easily.
- sparkie 12y agoIt's worth questioning why "proper use" of the file system is never done in practice. The filesystem is the wrong abstraction for this kind of problem, purely because it only offers one-to-many relationships, bar hacks like symbolic links which are rarely used where they'd be suitable, and because it's just simpler to dump metadata into a "container" with the media, and call it a file. Nearly all real-world data has many-to-many relationships - for example, in music, we have several artists for the same song, and of course, artists perform multiple songs. The container formats we use for music are really shoddy and try to force these organic relationships into a limited set of tags, and then our filesystem doesn't even have knowledge of this information unless we add plugins for specific file formats - another messy area. MusicBrainz and similar services organize music "the right way" in a RDBMS, and also includes support for things like multi-language tagging, which is sorely lacking in our filesystems. Instead of trying to come up with "one name to rule them all" for a piece of media, each object is just given an identity, a uuid, then all the metadata is related to the uuid. If we wanted to keep metadata and media separate, the obvious solution here is to dump the media in blobs in the same database - give each piece of media a uuid, and create a new relationship table to link the uuid to its metadata. I'd personally like this solution for my music. To do the same for files, we'd have an explosion of symbolic links to map the relationships, and no tool to really navigate through them effectively because we're missing a query language. If we sat these indexes on the filesystem, we'd have a big loss of performance because of the extra layers of indirection and searching, for which no optimization is done. A filesystem is really just a limited kind of database, but where we have a more powerful tool available, why not use it? Well, one reason is backward compatibility - our programs are written to look for files on filesystems, rather than streams from abstract sources. Perhaps what would be ideal in this kind of situation would be a FUSE layer which can expose the MusicBrainz database as a filesystem, but the underlying storage be the postgres db.
- tim333 12y ago>The filesystem is the wrong abstraction for this kind of problem From my experience having moved from PC to Mac, I find iPhoto trying to hide my files and abstract things away a pain in the arse. On the old computer I had them in folders labelled 2012, 2013 etc. I now want to move some previous years to another disk to save memory and it's hard to figure that now. Previously I would just have dragged a folder. I guess my point is that fancy database structures may be hard for the human brain to conceptualise.
- vezzy-fnord 12y agoWell, it's true that symlinks are a hack, but there is a proper solution to this problem that fits very neatly within the file system model: namespaces (http://www.cs.bell-labs.com/sys/doc/names.html http://www.cs.bell-labs.com/sys/doc/names.html).
- vidarh 12y ago> why even have a file system? A database seems more natural Fine, as long as that database is structured so I can arrange my own hierarchical structure of items, and manage them with something with the UI of a file manager. I've yet to come across any of these "treat my collection of X as a database" tools that have been satisfactory enough for me to maintain that collection only through that tool. Not one. That makes me doubt we're particularly close to doing away with more traditional databases. In terms of my own data, I've moved more and more away from databases. E.g. my blog, my personal wiki etc. used to be in databases, but are now plain files because it's far less painful to work with. Yes, I realise I'm not a regular user, but I've also seen enough "regular people" build deepl, complex hierarchies of files to realise that while some people may be satisfied with databases, many are not.
- gglitch 12y agoSame. Importing my mp3 collection into iTunes, for example, has been a disaster.
- TheSpiceIsLife 12y agoI often hear this criticism of iTunes, even from the tech savvy crowd, however there are two settings in iTunes: Preferences > Advanced (or whatever the Windows equivalent is) 'Keep iTunes Media folder organized' and 'Copy files to iTunes Media folder when adding to library'. Just make sure these are both off and iTunes won't touch your existing directory structure. With those settings off the iTunes is just another media player, with the convenience of streaming to a couple AirPort Express devices. But otherwise, yes, I agree, it makes a mess. I've forgotten to make sure these are off on a new machine once and a fresh install once, and rather than unmangling what iTunes does I just restored meda audio library from backup.
- seanmcdirmid 12y agoMost people aren't very good with filing cabinets, I mean, there are definitely people who are (the very organized ones), but the user data on an average desktop file system is pretty flat. Web directories failed for this reason, most people want to find stuff, not file stuff. I always thought that the right answer would be a system based on tags, which, if you allow for the tagging of tags, is fairly hierarchical. As a metaphor, the tag is put on data that exists in an otherwise flat unorganized space, and you would place tags on the data after the fact (now you create folders, and "move" data into those folders).
- informatimago 12y agoThe main problem is that of data ownership (which implies, having the data local, and processing it locally). Subsidiarily, having control over one own data sets, means that you can process them with various tools, and therefore that those data sets are independent from the tools. (I'm pointing the finger at you, iOS and Android). How this data is refered to, organized, and stored is irrelevant. The classic notions of file, hierachical directory, file systems are very useful and practical, because they promote full data set ownership and processability by user controled programs. Other kinds of systems may provide other notions (such as that of catalogs of objects, in capability based operating systems). But as long as the user has control over his own data, and can process it orthogonaly with his own programs, it doesn't matter much how it's stored. The problem indeed is when corporations lure users into giving up control of their data and processing with convenience. Users should be better educated. https://my.fsf.org/donate/ https://my.fsf.org/donate/ Oh, and developer should also try to provide convenience while preserving user ownership and control. (Unfortunately, it's very hard to do on platforms like iOS and Android, notably thru their app stores).
- motters 12y agoI think "data ownership" needs to be highlighted more prominently as a feature. I think it's a critical battle and that loss of ownership - as we can already see - leads to a lot of negative consequences for the user.
- sseveran 12y agoIn my experience users don't want to be educated, they want software that just works and is easy to use. They want transparent access to their documents wherever they are on any device without thinking about it. So to achieve this goal users want a piece of software (Drive,iCloud,Dropbox,etc...) to manage their data for them.
- wtbob 12y ago> The main problem is that of data ownership (which implies, having the data local, and processing it locally). I think that your concept of data ownership can be broken down into two related but independent concepts: access and control. Access means the ability to read; control means the ability to delete. Right now, if I upload content to almost any cloud provider I give access and control of my data to that provider: using G+ gives Google the ability to view every photo I upload, and to delete them all at a whim. In a hypothetical cloud provider which enabled me to encrypt data locally, they would have no access to my data, but would still be able to delete it. I would actually like the ability to grant control over a copy of my data without granting access to it; the two should be independent. I've had a vague idea for some time of system which enabled me to store encrypted data and share the keys with my friends but not with the storage provider. There's been some interesting work in indexing encrypted data which could be pertinent to this, enabling a storage provider to offer search capability while still being unable to read the data itself.