3 ms·
So this looks a lot like what Hadoop did with .har files[1] on HDFS (like storing GPS tiles on HDFS without blowing through the 1 M files per-dir limit). > It
by gopalv 5y ago
So this looks a lot like what Hadoop did with .har files[1] on HDFS (like storing GPS tiles on HDFS without blowing through the 1 M files per-dir limit).
> It is not possible to update individual files inside the ZIP file. Therefore this should only be used for data that isn’t expected to change.
I've actually done file-replaces on .zip files on HDFS, because .ZIP files are actually written with a directory in a footer, you can go to the end and append new data without having to "modify" existing files.
This doesn't conflict with the block level immutability, though the entire write has to be a single commit to avoid leaving the file in a bad way.
I'd say that the best case use-case for this is the storage of log files (like if you had fluentd writing .zip files by appending to it rather than a diff object for each 5 minute window).
When it comes to stuff which compresses well but full of small objects, ZIP is pretty bad because each file in the zip independently contains a dictionary (look at the .xlsx file inside to know how MSFT solved that, but in way which makes you hate it - a strings directory for shared strings across all files).
[1] - https://hadoop.apache.org/docs/r1.2.1/hadoop_archives.html#How+to+Create+an+Archive https://hadoop.apache.org/docs/r1.2.1/hadoop_archives.html#H...
- klauspost 5y ago> When it comes to stuff which compresses well but full of small objects, ZIP is pretty bad Correct, but it also allows you to independently access files, which is a win for this use-case. The goal isn't the compression itself, but reducing the number of objects, which in itself reduces the file system block overhead. Double checking, it seems like files compressed with zstandard method aren't supported. This will (when enabled) give both better compression and faster decompression. That should be added shortly.