4 ms·
Certain fields in the ZIP local file header require random access at write time. Although deflate is self-delimiting, for non-compressed items you must either k
by shockinglytrue 6y ago
Certain fields in the ZIP local file header require random access at write time. Although deflate is self-delimiting, for non-compressed items you must either know the size of item upfront, seek after learning the size, or force all readers to fully buffer the ZIP before decompressing any entry, in order to access the size stored in the central directory at the end of the ZIP
The CRC field of the local file header is similar, but its optionality is less painful than the length field
In other words, there isn't a single 'good' ZIP implementation, it all depends on what you're aiming for. AFAIK the built-in Python zipfile module always fully populates the local file header.
- tyingq 6y ago"Certain fields in the ZIP local file header require random access at write time...The CRC field..." I covered that in the sentence that described these as things you could fill with a placeholder and seek() back to later.
- shockinglytrue 6y ago> Kind of a shame someone has to specifically implement a low memory usage library. That implies other implementations went the lazy route. Streaming and seeking are antithetical, an implementation gets to pick one, and I imagine for most users streaming would be the edge case. It's debatable whether a library that ensures a fully spec-conforming output is generated at the expense of memory or IO is better or worse. (FWIW this is written from the perspective of someone who for perverse reasons once wrote a streaming ZIP reader. In order to build such a thing, the ZIP writer had to be buffering/seeking. You can't have both without giving up ability to store uncompressed assets without needless overhead (e.g. JPEG-in-deflate))
- tyingq 6y ago"Streaming and seeking are antithetical" I disagree, for a library creating zip files. It's the logical way to do it. Otherwise, you're slurping potentially big files into memory, or making un-needed copies. I've also written zip library code.
- shockinglytrue 6y agoWell, we're not discussing creating zip files where seeking is an option, we're discussing streaming in the context of the attached link which uses it concretely in terms of writing to a network.
- tyingq 6y agoThat explains our disagreement. I read "large zip archives" and assumed there was at least an intermediate file created.
- dependenttypes 6y agoIf you are streaming you should be using deflate directly instead.
- deleted 6y ago[deleted]
- zmodem 6y ago> Certain fields in the ZIP local file header require random access at write time. Isn't that what bit 3 in the general purpose bit flags if for? When that's set, the CRC and compressed/uncompressed file lengths are written in a Data Descriptor block after the member file data.
- shockinglytrue 6y agoIt's been a few years, but I seem to remember there simply was no way to detect end of non-compressed content lacking a length header without first reading the TOC except substring search. This would quickly get very messy when the non-compressed content might contain another embedded ZIP, etc.
- zmodem 6y agoThat's during extraction though. There's no way to reliably decompress a ZIP without reading the Central Directory first. While scanning for the Local File Headers works for most files, it's not guaranteed to be correct since it may not match what's in the Central Directory. Also as you point out it doesn't work when the file length is not in the LFH. Streaming zip creation is well supported by the format though.