3 ms·
They way I view is similar to de-normalization/"pre-materializing" in a database: the real solution is for a file to be tagged with pieces of metadata, maintain
by strlen 14y ago
They way I view is similar to de-normalization/"pre-materializing" in a database: the real solution is for a file to be tagged with pieces of metadata, maintain a real-time index, and then to be able to perform exact queries instantly (something like LinkedIn's faceted search) against this metadata.
E.g., first I want to find all files tagged with "code", then I want to find all files tagged with "current project". Then I get a list of all files in a current project and possible tags I could look for (e.g., it will show many are also tagged with "java" or with "C++" so I can add those tags to my search).
The problem is historically (that is, until recently) these kind of queries have been prohibitively expensive in both space and time: there's time complexity of doing a boolean query on the index, there's the space complexity of storing the index amenable to those queries both on memory and in-disk. This is difficult to in the context of legacy desktop machines (think minimal requirements to run Windows XP: 256mb of RAM, 5400 rpm disk, pentium III cpu) that were prevalent until just a few years ago.
Instead, the solution chosen is the same solution that folks building distributed "NoSQL" databases or sharded SQL database setup go for: group all related data into a few rows (in this case directories) and treat the data as hierarchical (a return to pre-RDBMS era).
Today if your working set fits in main memory (which is most probably the case for even the biggest data collections on most people's personal desktops), there's no longer a performance related reason to use a single-node (non-partitioned) "NoSQL"-style setup (there may be other reasons to, e.g., wanting to have a more flexible schema, saving time, wanting to scale out later, etc...)
I suspect the real reason we haven't done this with file systems is due to legacy: I have data organized into files and folders going back to ~1998 that I do not want to lose and can't afford to tag manually.
The support from my theory comes from the fact that the first systems to embrace non-hierarchical file system have been legacy-free devices where important storage (for 99% of the people out there...) is either on the SIM card (addresses), in the cloud (mail) or is de-facto tagged and has to be copied over in any case (photos which have EXIF tags and are stored in photo album apps or cloud services).
The same has also been happening on the other end of the spectrum: e.g., large scale distributed storage system for much similar reasons (i.e., they are the first of their kind, so there was a "license" to build these systems in a legacy-free fashion).
tl;dr Hierarchical storage is a non-functional requirement. Functional requirement is being able to do complex exact queries.