3 ms·
It would be interesting to try something like this on the file system level, yes. Basically you could heuristically find interesting files, then train a compre
by phiresky 4y ago
It would be interesting to try something like this on the file system level, yes.
Basically you could heuristically find interesting files, then train a compression dictionary for e.g. each file or each directory. Then you'd compress 32kB blocks of each file with the corresponding dictionary. Note that this would be pretty different from how existing file system compression works (by splitting everything into blocks of e.g. 128kB and just compressing those normally).
I think it would be hard to find heuristics that work well for this method. The result would be pretty different than what I'm doing here with different tradeoffs. sqlite-zstd is able to use application-level information of which data is similar and will compress well together. You couldn't do that within the filesystem. It also works "in cooperation" with the file format of the database instead of against it - e.g. you probably wouldn't want to compress the inner B-tree nodes.
On the other hand, it would then work for any file format not just SQLite databases. E.g. a huge folder of small json files