11 ms·
I've dealt with a different situation before and handled it with sharding. It's pretty easy to retrofit in code, and it's not hard to do a one-time move of exis
by gregmac 4y ago
I've dealt with a different situation before and handled it with sharding. It's pretty easy to retrofit in code, and it's not hard to do a one-time move of existing files to the new layout.
Simply put, you make an algorithm like:
let hash = md5(filename)
let folderPath = path.join(root, hash.substr(0,1), hash.substr(0,2), hash.substr(0,3), filename)
The original name can be used to derive the sharded path, so you don't need a lookup database or anything crazy like that.
You end up storing files like:
.../7/7f/7fa/somefile.png
This scales to whatever depth you want: with 3 levels and 1.2 million files, you can expect about 1200000/16^3 = 293 files per directory, which is trivial.
There's lots of other strategies, too, depending on your needs:
# filename are already GUIDs or hashes:
.../1/1f/1fd/1fd2dd27-d307-47b2-b37d-11903fd0f03d.png
# date-based to make cleaning up old files easy:
.../2023/01/19/app.log
# hash by existing numeric identifier (like customerid) using modulus:
.../customers/4/94/52194/customerfile.dat
- pc86 4y agoThis seems like a better solution that just tweaking Linux to make it take only "a few seconds" to pull up an image. I'm curious what downsides/traps there are in this approach vs. sticking with the single flat directory.
- bluedino 4y agoI have used that method many times with success. One side-effect is you can't go into a folder and type something like: `cp .doc` or `cp screenshots xxxxx`
- thesneakin 4y ago./**/*.doc Or `find -exec`
- bluedino 4y agoExcept that will be very slow when you have millions of files
- c0l0 4y agoEnumerating all images in the system over the network is what took a few seconds. "Pulling up an image" worked nigh-instantly, because contemporary filesystems have (and had) efficient indexes for the file(name)->data mapping(s) involved.
- NovemberWhiskey 4y agoRight; you're basically using the filesystem to implement an index based on the attributes of the file according to the expected usage pattern.
- layer8 4y agoI wonder what the sweet spot is performance-wise regarding hierarchical depth vs. the number of entries in each directory, between the two extremes of putting all files in a single directory (depth 0) and using only one bit per directory (depth log(number of files)).
- thrashh 4y agoDepends on the filesystem that you use, like Dynamo vs Postgres OP’s post probably involved an old filesystem. For example, ext2 did not index the files within a directory, but ext3 introduced H-trees which made big directories much more feasible.
- layer8 4y agoOf course it depends on the filesystem. I’d be interested in what it is concretely, for each of the commonly used desktop filesystems.
- c0l0 4y agoYes sure, the "successor" system that was eventually implemented to redundantly host BLOBs of all kinds used a very similar scheme - it's a well-known and effective technique for, e.g., on-disk caching in popular HTTP server implementations. The problem that kept us from doing something like this with out pile of 1.2M images was that there were several different codebases (and even a hand full of remote sites which had code that we knew relied on this assumption, and that we could not directly affect or change) that tied into this flat directory full of image junk in various ways (via the local filesystem, NFS, over HTTP, and I think even FTP was in use for a while after I had originally joined), which made it very hard to push through any breaking changes, no matter how architecturally sound they would have been.
- nine_k 4y agoThe beauty of this is that you can slowly migrate data to the new schema, while serving files. The updated code first looks up a new path (pretty fast), if it's not found, it looks up the old path. Only the original file name is required, and the code change is literally 10-20 lines, including the function that derives the new path.