8 ms·
Why Fred Wilson is wrong – files aren’t dead
- whyleyc 12y agoAuthor here - would love to hear your comments, as I think this is an interesting area of debate, especially with digital obsolescence waiting to bite us all.
- sliken 12y agoWell a file is just a special case of an object, one with a single parent (a directory), and a name typically designed to be human readable. One that's particularly limiting metadata. Typically owner, name, and a few time stamps. So sure people want to download streaming videos. But chances are the result will be put into something that lets them view them by director, title, genre, newest, etc. Generally it seems more natural and user friendly to have your photos, documents, music, and videos indexed. That way you don't have to remember the name, directory, or even what you computer you used to access it. You just look it up by when you used it last (i.e. I edited it yesteday) or by whatever metadata you remember. So if a user's main method of access is by index of metadata not available in the filesytem, why even have a file system? A database seems more natural. After all why is /directory/fi lename more important than being able to look based on arbitrary metadata?
- altcognito 12y ago> why even have a file system? None of this implies that the underlying file systems aren't really useful in this situation. In fact, proper use of a file system would include separation of the media indexes from the media themselves, so that if you want to migrate to a better index, you can do so easily.
- sparkie 12y agoIt's worth questioning why "proper use" of the file system is never done in practice. The filesystem is the wrong abstraction for this kind of problem, purely because it only offers one-to-many relationships, bar hacks like symbolic links which are rarely used where they'd be suitable, and because it's just simpler to dump metadata into a "container" with the media, and call it a file. Nearly all real-world data has many-to-many relationships - for example, in music, we have several artists for the same song, and of course, artists perform multiple songs. The container formats we use for music are really shoddy and try to force these organic relationships into a limited set of tags, and then our filesystem doesn't even have knowledge of this information unless we add plugins for specific file formats - another messy area. MusicBrainz and similar services organize music "the right way" in a RDBMS, and also includes support for things like multi-language tagging, which is sorely lacking in our filesystems. Instead of trying to come up with "one name to rule them all" for a piece of media, each object is just given an identity, a uuid, then all the metadata is related to the uuid. If we wanted to keep metadata and media separate, the obvious solution here is to dump the media in blobs in the same database - give each piece of media a uuid, and create a new relationship table to link the uuid to its metadata. I'd personally like this solution for my music. To do the same for files, we'd have an explosion of symbolic links to map the relationships, and no tool to really navigate through them effectively because we're missing a query language. If we sat these indexes on the filesystem, we'd have a big loss of performance because of the extra layers of indirection and searching, for which no optimization is done. A filesystem is really just a limited kind of database, but where we have a more powerful tool available, why not use it? Well, one reason is backward compatibility - our programs are written to look for files on filesystems, rather than streams from abstract sources. Perhaps what would be ideal in this kind of situation would be a FUSE layer which can expose the MusicBrainz database as a filesystem, but the underlying storage be the postgres db.
- tim333 12y ago>The filesystem is the wrong abstraction for this kind of problem From my experience having moved from PC to Mac, I find iPhoto trying to hide my files and abstract things away a pain in the arse. On the old computer I had them in folders labelled 2012, 2013 etc. I now want to move some previous years to another disk to save memory and it's hard to figure that now. Previously I would just have dragged a folder. I guess my point is that fancy database structures may be hard for the human brain to conceptualise.
- vezzy-fnord 12y agoWell, it's true that symlinks are a hack, but there is a proper solution to this problem that fits very neatly within the file system model: namespaces (http://www.cs.bell-labs.com/sys/doc/names.html http://www.cs.bell-labs.com/sys/doc/names.html).
- vidarh 12y ago> why even have a file system? A database seems more natural Fine, as long as that database is structured so I can arrange my own hierarchical structure of items, and manage them with something with the UI of a file manager. I've yet to come across any of these "treat my collection of X as a database" tools that have been satisfactory enough for me to maintain that collection only through that tool. Not one. That makes me doubt we're particularly close to doing away with more traditional databases. In terms of my own data, I've moved more and more away from databases. E.g. my blog, my personal wiki etc. used to be in databases, but are now plain files because it's far less painful to work with. Yes, I realise I'm not a regular user, but I've also seen enough "regular people" build deepl, complex hierarchies of files to realise that while some people may be satisfied with databases, many are not.
- gglitch 12y agoSame. Importing my mp3 collection into iTunes, for example, has been a disaster.
- TheSpiceIsLife 12y agoI often hear this criticism of iTunes, even from the tech savvy crowd, however there are two settings in iTunes: Preferences > Advanced (or whatever the Windows equivalent is) 'Keep iTunes Media folder organized' and 'Copy files to iTunes Media folder when adding to library'. Just make sure these are both off and iTunes won't touch your existing directory structure. With those settings off the iTunes is just another media player, with the convenience of streaming to a couple AirPort Express devices. But otherwise, yes, I agree, it makes a mess. I've forgotten to make sure these are off on a new machine once and a fresh install once, and rather than unmangling what iTunes does I just restored meda audio library from backup.
- seanmcdirmid 12y agoMost people aren't very good with filing cabinets, I mean, there are definitely people who are (the very organized ones), but the user data on an average desktop file system is pretty flat. Web directories failed for this reason, most people want to find stuff, not file stuff. I always thought that the right answer would be a system based on tags, which, if you allow for the tagging of tags, is fairly hierarchical. As a metaphor, the tag is put on data that exists in an otherwise flat unorganized space, and you would place tags on the data after the fact (now you create folders, and "move" data into those folders).
- informatimago 12y agoThe main problem is that of data ownership (which implies, having the data local, and processing it locally). Subsidiarily, having control over one own data sets, means that you can process them with various tools, and therefore that those data sets are independent from the tools. (I'm pointing the finger at you, iOS and Android). How this data is refered to, organized, and stored is irrelevant. The classic notions of file, hierachical directory, file systems are very useful and practical, because they promote full data set ownership and processability by user controled programs. Other kinds of systems may provide other notions (such as that of catalogs of objects, in capability based operating systems). But as long as the user has control over his own data, and can process it orthogonaly with his own programs, it doesn't matter much how it's stored. The problem indeed is when corporations lure users into giving up control of their data and processing with convenience. Users should be better educated. https://my.fsf.org/donate/ https://my.fsf.org/donate/ Oh, and developer should also try to provide convenience while preserving user ownership and control. (Unfortunately, it's very hard to do on platforms like iOS and Android, notably thru their app stores).
- motters 12y agoI think "data ownership" needs to be highlighted more prominently as a feature. I think it's a critical battle and that loss of ownership - as we can already see - leads to a lot of negative consequences for the user.
- sseveran 12y agoIn my experience users don't want to be educated, they want software that just works and is easy to use. They want transparent access to their documents wherever they are on any device without thinking about it. So to achieve this goal users want a piece of software (Drive,iCloud,Dropbox,etc...) to manage their data for them.
- wtbob 12y ago> The main problem is that of data ownership (which implies, having the data local, and processing it locally). I think that your concept of data ownership can be broken down into two related but independent concepts: access and control. Access means the ability to read; control means the ability to delete. Right now, if I upload content to almost any cloud provider I give access and control of my data to that provider: using G+ gives Google the ability to view every photo I upload, and to delete them all at a whim. In a hypothetical cloud provider which enabled me to encrypt data locally, they would have no access to my data, but would still be able to delete it. I would actually like the ability to grant control over a copy of my data without granting access to it; the two should be independent. I've had a vague idea for some time of system which enabled me to store encrypted data and share the keys with my friends but not with the storage provider. There's been some interesting work in indexing encrypted data which could be pertinent to this, enabling a storage provider to offer search capability while still being unable to read the data itself.
- jacquesm 12y agoThe commoditization (sp?) of your data is what the walled gardens are all about. So it's in the interest of companies like Apple, Google, Yahoo, Dropbox, Microsoft and so on to remove as many of your files from your control as they can get away and to package them in a non-file metaphor so they can sell you your own data back, either directly through access fees or indirectly by selling you advertising right along your own data. I wished more people would see the endgame in situations like these: you're going to be paying through the nose for something that was already yours. Hosting your own data is trivial, the only case where I can see your data moving to the cloud is for backup purposes, off-site is better than on-site in case of disaster recovery. So, have a good and long look at that firewall that protects you from the big bad cyber terrorists out there. That same firewall that stops the bad guys from coming in (and your ISP by blocking access to port 80 and a couple of others) are what keeps the peer-to-peer potential of the web from being realized. So when VCs start trumpeting the 'end of files' make sure you realize what you're giving up.
- unreal37 12y agoNot worried about them selling me my own data, so much as they want to get rid of the concept of "buy once and own it forever". The Internet megacorps have become addicted to the idea of lifetime rental. Rent MS Office. Rent this movie. Rent this song. They're getting rid of the concept of software and media ownership for the sake of profit.
- jacquesm 12y agoRent your own photos. Rent the software you use. It's all part and parcel of the same concept. A subscription model rather than an ownership model. The iPod was a trojan horse.
- arrrg 12y agoSeeing as Apple in particular has to be dragged into this new world and everyone (really everyone) was there before them I don’t think there is any truth at all to that characterisation. The iPod was never intended to be a trojan horse. That’s revisionist history at its best.
- unreal37 12y agoThe author starts by tearing down a straw-man. Fred Wilson wasn't predicting the end of "file systems". He was saying people go to the cloud to consume content (Netflix, Youtube, Spotify) and go to the cloud to create content (Google Drive, Office 360, Soundcloud, etc). And they don't need to concern themselves with files. Fred Wilson is a smart guy, and knows that even iOS, Android, etc run on file systems. I disagree files are dead too. But file systems is clearly not what he was talking about.
- davidw 12y agoHe does go on to admit this, more or less, in the article. So Fred Wilson remains right, and there's not much to see here, IMO.
- unreal37 12y agoI agree. But the headline is "Why Fred Wilson is wrong". And the main argument is not something he said.
- whyleyc 12y agoAuthor here - I tried to keep "on topic" and constrain myself to files, but you're right I did drift into filesystems too which probably wasn't helpful. I think the core argument still holds though - which is that some people (Fred included ?) see files dying out as inevitable, whereas I'm not sure that's true (for the various reasons I cite).
- uptown 12y agoI found it odd that he used Dropbox as the argument for the death of files. While Dropbox does enable apps to sync data relatively seamlessly, the vast majority of Dropbox users I've encountered use it to store individual files or folders of individual files.
- jacquesm 12y agoI wrote this a while ago, and it's too long for a comment here anyway: http://jacquesmattheij.com/the+dropbox+endgame http://jacquesmattheij.com/the+dropbox+endgame
- uptown 12y agoCertainly possible. Comes down to how much control content creators are willing to give-up, and what partnerships Dropbox forges with software developers. For example, will Adobe want a user's entire photo library synced with Dropbox? What's in it for them? How do they handle the transitional period where some users want that and others don't?
- scholia 12y agoIf you buy one of those $69-$99 Windows 8.1 tablets with Bing, you'll find some programs save to OneDrive by default. Once you've logged in, it's basically a transparent part of the file system.
- teach 12y agoI think Google was the first to deliver on this promise; Chromebooks have no user-accessible local storage and all the content is in Google Drive. I guess it turns out to have been cheaper to build their own almost-as-good Dropbox replacement than to partner with them.
- kuschku 12y agoI’m actually using MEGA like this, I sync my home folder to MEGA and all devices. Really useful.
- tim333 12y ago
- toksaitov 12y agoI think the author is missing the point of the original article. The point was not that files are going away. It's just that the concept of file is not that important for the majority in 2014. Non-tech people go to Spotify to listen to music, open Netflix to watch a movie, work with documents on Google Drive and our Grannies don't have to learn what a file is to use their phones (for something more than just calling us).
- zaphar 12y agoHe didn't miss the point. He says exactly that later in the article. He just expands on the distinction a bit is all.
- bitL 12y agoI can't imagine getting rid of files and going all cloud in my creative visual art projects - just one second of a 14-bit RAW YUV 4:4:4 5k movie takes gigabytes, even the latest M.2 PCIe SSDs have trouble playing it back realtime, not mentioning "slow" 1Gbit Internet connection at best, though definitely not for upload. I have a feeling we are getting backwards from the efficiency point of view and am thinking this won't be a sustainable way forward - even in distributed/parallel algorithms you try to keep locality to improve performance and save resources, not to transfer everything back and forth via some middle man just because some business guy came up with some "genial" idea how to milk money. Internet companies have a few chances to earn money, subscription being one of them, but massive push into useless cloud for their particular business cases like in the case of Adobe CC or Microsoft Office just leaves bad aftertaste. Also, Dropbox/GDrive don't allow incremental update API calls for 3rd party apps, which wastes precious upload bandwidth, and in addition end to end encryption fully controlled by user is not provided which would justify full uploads. All of this just screams of artificial constraints which do not benefit anyone, in the long term not even to those companies.
- k-mcgrady 12y ago>> "I can't imagine getting rid of files and going all cloud in my creative visual art projects - just one second of a 14-bit RAW YUV 4:4:4 5k movie takes gigabytes, even the latest M.2 PCIe SSDs have trouble playing it back realtime" How many people have this or a similar problem though? I would guess that most people deal with word documents, jpegs, and mp3's on a daily basis and rarely encounter much else.
- wslh 12y agoThe "bitL problem" is part of a big industry, so a lot of people/companies have this problem. We can say that ~100% of the filmmakers should deal with this all the time. Probably this is why YC funded companies such as http://www.wireover.com/ http://www.wireover.com/
- bitL 12y agoUntil transfer over Internet is faster than sending a harddrive via FedEx, there is not much hope. You need 4k+ camera for acceptable quality 1080p final result, 16k+ camera for good 4k, 32k+ for 8k, and this just brings to their knees the most powerful Intel processors or storage devices, not mentioning network... :-(
- mark_l_watson 12y agoAs people have already pointed out, Fred Wilson is mostly talking about an abstraction layer so most users think about applications, which provide an interface layer to files they use. I think that devices with a relatively small amount of SSD storage contribute to cloud and web service providers getting more control of people's data. Selective sync in services like OneCloud, Google Drive, and Dropbox allow users to just keep local copies of files they need in the near future. At least I do this.
- mark_l_watson 12y agoAs people have already pointed out, Fred Wilson is mostly talking about an abstraction layer so most users think about applications, which provide an interface layer to files they use. I think that devices with a relatively small amount of SSD storage contribute to cloud and web service providers getting more control of people's data. Selective sync in services like OneCloud, Google Drive, and Dropbox allow users to just keep local copies of files they need in the near future. At least I do this.
- deleted 12y ago[deleted]
- deleted 12y ago[deleted]
- Kiro 12y agoThe upward trends for downloading SoundCloud, Netflix and YouTube stuff are most likely due to their rise in popularity so not sure if it's relevant at all.