4 ms·
If the customers stored .gz/.zip/... files you could transparently transcompress them to .zstd and back (with an added test of course)
by hyuijk 4y ago
If the customers stored .gz/.zip/... files you could transparently transcompress them to .zstd and back (with an added test of course)
- sgtnoodle 4y agoIt seems like that would cause problems. Right off the bat, any integrity hashes like md5 or sha256 for the original compressed files would likely be corrupted. Also, the compressed archive could have been structurally baked in a specific way that's meaningful to the customer. Zip archives in particular can have arbitrary data pretended to them. I suppose you could speculatively decompress and then re-compress and see if you get the original compressed file back, and maybe most people happen to use the same compression implementations with default settings.
- mike_hock 4y agoThat wastes a lot of CPU compared to just running it through zstd.
- spockz 4y agoYou would only apply the compression on the internally stored file and then decompress when retrieving it for the customer again. That way all the hashes and original structure of the user are retained.
- klauspost 4y agoYou would need to be able to reconstruct the input file bit-by-bit. S3 Clients expect to get back what they sent, exactly. This puts a serious limitation on your compression. You would only be able to re-do the entropy coding part of DEFLATE, which is actually pretty good. You would still need to store the original Huffman tables for each block, so you can reconstruct the entropy coding exactly. I doubt this would even gain you a single percentage.
- mkup 4y agoBesides ZIP metadata, there may be flushes in specific points in the deflate stream (which applies to .gz files as well). These flushes reset the compression dictionary and make further compressed data independent from previous data (at the expense of losing some compression efficiency). So: AWS S3 customer may have injected these flushes to their .gz files (for whatever reason, e.g. steganography), and after gzip-to-zstd-to-gzip transcompression this steganographic data will be be lost (and of course sha256 and other similar hashes will be different, as you already said).
- sgtnoodle 4y agoExactly. I've built several logging systems over the years that intentionally flush compression streams (or concatenate gz streams) for robustness reasons.
- lifthrasiir 4y agoThere are tools like preflate [1] or precomp [2] that guarantees a bitwise identical reconstruction, of course modulo bugs. [1] https://github.com/deus-libri/preflate https://github.com/deus-libri/preflate [2] https://github.com/schnaader/precomp-cpp/ https://github.com/schnaader/precomp-cpp/ (which internally makes use of preflate)
- Cyberdog 4y ago> Also, the compressed archive could have been structurally baked in a specific way that's meaningful to the customer. EPUB is an example of this. They're mostly bog-standard ZIP archives, but in order for the file to be valid, the "first" file in the archive, linearly speaking, must be named "mimetype" and stored with no compression (each file in a ZIP archive can have a different compression level). If you just unzip an EPUB and then just dumbly zip it back up again, your end file will not be a valid EPUB.