5 ms·
On my AWS service (EFS) many customers do use lots of small files. And for legacy apps there is nothing much you can do, but as a PSA, use the right tool for th
by harshaw 3y ago
On my AWS service (EFS) many customers do use lots of small files. And for legacy apps there is nothing much you can do, but as a PSA, use the right tool for the job :) many people use file systems as databases when a database would be more appropriate.
Put it this way - moving a sqlite file with 1 million rows is much easier than moving 1 million files.
- svara 3y agoWhy though? If you squint a little, a filesystem and a database are the same thing and filesystem semantics are useful and well understood. Performance is "just" an engineering issue.
- edgyquant 3y agoThey are not the same thing. In general Databases allow for querying of the structured contents inside of a “file” but also they tend to exist (now days at least) as an abstraction on top of the file system.
- metalliqaz 3y agowell you're very correct that a filesystem is just a kind of database. So the equivalent of moving the sqlite database would be to simply move the filesystem image rather than each file individually. However I think the issue is that the OS and services all run on top of the filesystem, so doing operations at the level of block storage is not practical. But... if the system is virtulized, I guess you could make it work
- 01HNNWZ0MV43FF 3y agoWith LVM you can do block device snapshots with virtualizing the whole system, but I don't think there's any way to get around the fact that, e.g. a 1 TB filesystem with 600 GB of files will have 400 GB of garbage that you end up copying, whereas a SQLite or Postgres DB would vacuum itself periodically and actually take up only 600 GB.
- 01HNNWZ0MV43FF 3y agowithout virtualizing*
- still_grokking 3y agoDepends on the FS and tools used. With for example XFS and `xfsdump` / `xfsrestore`, or ZFS and `send` / `receive` you would only copy the data, and something like a "vacuum" would happen also.
- harshaw 3y agosay you want to store 10 million records. if you do this as a files in a say, one big directory (or even a sharded directory hiearchy), you have to update the parent directory (inode) when you create a new file (because the directory mtime needs to change) That's extra overhead vs what you can get with appending to a journaled database. POSIX encodes these specifications of what a file system is - so there isn't much ability to innovate here. And yes - we work really hard to address the performance issues here, but there are constraints.
- charcircuit 3y agoAbandon POSIX and use the filesystem over userspace
- marwis 3y agoBut why can't journaled filesystem do the same? Didn't Log Structured Filesystem work like that? Or SQLite over Fuse.
- duped 3y agoAre database semantics not useful or well understood? There's also no ways to really make a file system faster when dealing with many small files and directories. If you can avoid doing it by using a database you probably should.
- Asraelite 3y ago> Are database semantics not useful or well understood? Sort of, yeah. Consider how many command line programs will accept a file as an input parameter vs. how many will accept an SQL query. Existing tooling is predominantly designed to work with files so depending on what you're trying to accomplish, you may have a lot more work to do when using a database.
- MichaelZuo 3y agoThat’s an interesting point, in terms of shrink wrapped software its easily 100x more software that will accept a file input.
- qubex 3y agoWell yes, but also it varies… on IBM AS/400 machines running OS/400 (currently iSeries if I’m not mistaken) the file system is basically a database and commands basically parse queries.
- vineyardmike 3y agoAnd as well all know, an IBM AS/400 is the most common computer for serious software development in 2024.
- qubex 3y agoWhy bother staying silent when one can slide some sarcasm in, eh?
- declaredapple 3y ago> There's also no ways to really make a file system faster when dealing with many small files and directories. Depends on the file/directory structure. Traditionally back in the CGI days URI's were links to files on a server. They often still are for static files. If you know the path to a file then the lookup should be very fast. But yeah any other type of index other then the path isn't really possible.
- wongarsu 3y agoDatabases are optimized for storing a couple trillion records of a couple dozen bytes each. File systems are optimized for storing a couple million records of a couple Megabytes to Terabytes each, with maybe a hundred bytes or so of metadata attached. They make very different engineering tradeoffs. And while you can invest time to make them better at each other's tasks, there is only so much you can do.
- MichaelZuo 3y agoIsn’t that what ZFS is trying to solve? Such that storing a couple billion “ records of a couple Megabytes to Terabytes each” become feasible?
- deleted 3y ago[deleted]
- fluoridation 3y agoNo, ZFS tries to solve the problem of data integrity mainly. It also supports very large volumes and files, but it's not particularly good at storing billions of files in a single one. In that respect it's like any other file system. It'll do it, but you need to design your directory structure well to have fast access, and if you try to copy the directory structure it'll still be very slow.
- MarkSweep 3y agoEven in ZFS, there is more overhead for a file than I would guess a database has for a row. Each file needs a DNode, which is 512 bytes. It also needs an entry in it's parent directory's list of files (a 64 byte micro ZAP entry for example).
- rubiquity 3y ago> If you squint a little, a file system and a database are the same thing This is a tale and argument as old as time. It’s the squinting that is the problem. The devil is in the details.
- qubex 3y agoBeOS squinted so hard its native file system BeFS came out as an impressionist painting: the live metadata made it very database-like in behaviour with live search results (back in the mid nineties) but fairly lacklustre “actual file system” credentials.
- michael1999 3y agoTracking atime, atime, and full permissions on every dirent is a lot of work. Most people putting a bunch of small files on disk don't want any of that, and usually misunderstand the actual transaction promises of their filesystem. They would be better off with SQLite and a single table (name, blob).
- jazzyjackson 3y agohow about a ZFS snapshot of a million files, is that any different than a single file database?
- nyrikki 3y agoUnless you need filesystem features, object stores are probably a better target. Or possibly a document based data store. The monolithic persistence layer is an anti pattern unless you have specific reasons to support it. ZFS is quite nice as filesystems go but it is still a filesystem.
- bgm1975 3y agoAs the saying goes: If the only tool you have is a hammer, it is tempting to treat everything as if it were a nail.