3 ms·
The problem with this is that if your records/documents are small, you're wasting huge amounts of space because each file uses a full filesystem block. If you h
by ludocode 5y ago
The problem with this is that if your records/documents are small, you're wasting huge amounts of space because each file uses a full filesystem block. If you have, say, ten thousand records where each is 200 bytes, a decent database would store that in a bit over 2MB. Storing these as individual files on a filesystem with 4kB blocks will take up at least 40MB. This is a huge amount of wasted space, not to mention slow. (Some filesystems do support tail packing but that won't fully solve the problem.)
Not to mention all the other problems with this. The filesystem has a complete lack of higher-level features: no transactions, no snapshots, no indexing beyond filenames, no easy robustness guarantees (doing fsync() properly is a lot more complicated than it appears.) Honestly for modern apps the filesystem is just terrible at storing any internal mutable app data.
Once you start writing code to store auxiliary indices, synchronize writes, or pack multiple records per file, well at that point you're just implementing your own database. This might make sense if, say, you have a special way of compressing your data (like git). But generally you're better off using a real embedded database.
- inshadows 5y agoSome filesystem can store small data directly in inode. Limit for ext4 is 160 bytes though. https://unix.stackexchange.com/questions/197570/is-it-possible-to-store-data-directly-inside-an-inode-on-a-unix-linux-filesyst https://unix.stackexchange.com/questions/197570/is-it-possib... https://ext4.wiki.kernel.org/index.php/Ext4_Disk_Layout#Inline_Data https://ext4.wiki.kernel.org/index.php/Ext4_Disk_Layout#Inli...
- deleted 5y ago[deleted]
- rektide 5y agobtrfs has this too. it'd be cool if dirent_t could those this so one could quickly iterate thought these inline data. https://www.gnu.org/software/libc/manual/html_node/Directory-Entries.html https://www.gnu.org/software/libc/manual/html_node/Directory...
- fulafel 5y agoYou'll only be wasting space proportional to the number of objects, and the overhead is much smaller on ext4 (default fs on most linux distros) like the sibling comment explains. Most databases are quite small, and most of the rest are less than huge, so in most cases you won't be wasting huge amounts of space.
- quantumofalpha 5y agoOverhead seems to be about 4KiB, at least using default ext4 parameters: # zfs create -V 100G -b 4096 tank/test && mkfs.ext4 -v /dev/zvol/tank/test && mount /dev/zvol/tank/test /mnt && cd /mnt # df -k . Filesystem 1K-blocks Used Available Use% Mounted on /dev/zd16 102626232 24 97366944 1% /mnt # for x in `seq 1000000`; do echo $x >$x; done # create 1M tiny files # df -k . Filesystem 1K-blocks Used Available Use% Mounted on /dev/zd16 102626232 4022348 93344620 5% /mnt # bc -l (97366944-93344620)*1024/1000000 # free space diff per file 4118.859776 Plus you're wasting an inode per record - a limited resource in ext4, increasing which requires reformatting. You'd probably run out of inodes much sooner than out of space.
- fulafel 5y agoInteresting, indeed the inline data option is not enabled even by the latest e2fsprogs even though the feay has been there a long time. Re inodes, this is a good point too. These definitely reduce the size of db that fs works nicely for.
- brundolf 5y agoFor simple stuff I've had success keeping an in-memory data structure as the source of truth, and just persisting the whole thing to a file (JSON or otherwise) via a debounced function. Assuming you only have one process (or at least one main process), you only have to read the file on startup and can be pretty relaxed about your write strategy