3 ms·
When I read "If the input X is incompressible, then a copy of X is returned", I worry that this is broken. If I archive a file, then extract it from the archive
by reacweb 5y ago
When I read "If the input X is incompressible, then a copy of X is returned", I worry that this is broken. If I archive a file, then extract it from the archive, I can not be sure to obtain the same file. If the file is already compressed at the beginning, it will be decompressed at the end.
Maybe I am wrong. I didn't know this tool. My brief review of the documentation leads me to believe that it has an obvious problem.
- kybernetikos 5y agoI think your concern is dealt with in the immediately following section: >The Y parameter is the compressed content (the output from a prior call to sqlar_compress()) and SZ is the original uncompressed size of the input X that generated Y. If SZ is less than or equal to the size of Y, that indicates that no compression occurred, and so sqlar_uncompress(Y,SZ) returns a copy of Y. The format stores both the possibly compressed blob and the original size. From those two pieces of information it can always return the correct original file.
- reacweb 5y agoOk, I was wrong. My review was too brief.
- mbreese 5y ago> If SZ is less than or equal to the size of Y, that indicates that no compression occurred, and so sqlar_uncompress(Y,SZ) returns a copy of Y. This is pretty common for compression tools. If the input is incompressible (or not compressible by X%), then the original data is stored. In this case, the code is checking the stored size against the uncompressed size. If they are equal, then uncompress is a noop. My take away is that there is a zlib compression function built into SQLite. Which can be pretty handy. Another benefit I can see is that because the SQLite database has a flexible schema, you could add new features to the archive while maintaining backwards compatibility. For example, if you wanted to add a SHA1 hash to each record, you should be able to, while still allowing older tools to read the updated file.
- formerly_proven 5y agoIt's still a really annoying design, one should really explicitly communicate if and what compression method was used for a given blob of data. Here another DB column would have been a very easy way.
- lifthrasiir 5y agoNote that this kind of fallback is also prevalent in compression stream formats including zlib which sqlar uses, so the archive format doesn't need to reimplement the same fallback. The only reason it might be useful is the opportunistic support for random access for uncompressed data.