5 ms·
Python – Writing large ZIP archives without memory inflation
- chrisseaton 5y agoDoes writing ZIP archives usually cause memory inflation? I thought the algorithm had a simple fixed-size sliding window?
- TonyTrapp 5y agoI was thinking the same. Unless you're doing something really wrong, or if you are intentionally writing to a memory buffer, memory consumption of ZIP compression should be minimal across the whole process.
- stefan_ 5y agoYou would think that but it is Python so the likelihood is high that the magic convenient one liner is doing the equivalent of reading the file into a huge buffer and compressing that huge buffer at once.
- userbinator 5y agoI've seen this (antipattern?) in a lot of other HLLs too. People just don't seem to have a good grasp of the difference between the filesystem and memory, especially if all they've ever used and started with are HLLs.
- lifthrasiir 5y agoYou do need to keep the central directory record to be written at the very end, so it might cause a concern if you have a lot of (probably small) files. But Zipfly doesn't seem to solve that anyway.
- AshamedCaptain 5y agoSo on one top-level comment we have someone arguing about just writing the entire file to disk and don't bothering; while on the other top-level comment someone argues that the zip algorithm should already be using constant memory anyway. Funny to see the usual great programming school divide so succinctly put :) Anyway, already seen a number of times https://news.ycombinator.com/from?site=github.com/buzonio https://news.ycombinator.com/from?site=github.com/buzonio
- rndgermandude 5y agoZIP in itself can have different compression algorithms. deflate and "no compression" (store) being the most common by far (because compatibility reasons), followed by deflate64 and and a long tail of things like bzip2, zstd, lzma and others (the ones I mentioned are some of the methods "officially" assigned a method number by pkware). A lot of software uses zlib to implement zip, and there the only real options for broad compatibility are therefore deflate and no compression. deflate has a small fixed-size window buffer, with no compression you can avoid such a buffer altogether. No compression is quite alright if you put data into the archive that cannot be compressed efficiently anyway, like already compressed multi-media formats, and it will save you and whoever uses the archive later a bunch of compute. The only thing that may eat up memory is the final name directory in the "footer". No compression also has the added benefit that you can pre-compute the final size of the zip file if you know what goes into the zip already (it's just a matter of summing up the sizes, and the sizes of the info entries and final footer entries), which can be nice when sending the result over http, as it enables you to send a Content-Length, instead of not sending one because you don't know it yet, using chunked transfer encoding and letting the browser show some "20MB of unknown downloaded". Then it really depends on the implementation. A lot of libraries will buffer in an entire input file, compress it in one go and store the result back to disk or some buffer, and things like that. But it is perfectly possible to implement a zip writer that implements a stream/callback interface to read the result compatible with whatever your language is using (e.g. C++ istream (yeahyeah :P), C# Stream, python file-like objects, node Readable streams, etc). With the python standard library zipfile implementation, while it is possible to open "file-like objects" (i.e. objects implementing read/write/seek, like a BytesIO buffer), it isn't really possible easily to drain the data after just a chunk of it is done (or have it output to network/http streams, as these usually do not or cannot implement seek). The OP implementation solves this by defining everything you want in the zip beforehand and then using python generators to get the data out in chunks that you can then write to whereever you want sequentially.
- herpderperator 5y agoI don't think they literally mean "for immediate sending out to clients" right? Because that would block the worker thread for the duration of the download. It makes more sense to write the output to disk and serve it out as normal static content.
- deleted 5y ago[deleted]
- thrashh 5y agoIt would only block if you did it synchronously. If I had to build an endpoint that generated ZIPs on the fly, I would explicitly build it to be asynchronous and efficient as possible. The only issue is that you might want to attach a buffer somewhere that might dump to disk past a certain amount, if the source of your data can’t be “paused.”
- rodmena 5y agoNice. The code quality could be better though.
- deleted 5y ago[deleted]
- waydegg 5y agoFYI the link to the help page (on your website) doesn't return anything (https://buzon.io/en/help/ https://buzon.io/en/help/).
- scosman 5y agoI built a similar project in Go, along with http server. Comes in handy for streaming (as you point out): https://github.com/scosman/zipstreamer https://github.com/scosman/zipstreamer
- simonw 5y agoThis is really interesting. My https://datasette.io/ https://datasette.io/ application offers features to export relational data, using Python asyncio under the hood. It can currently stream an arbitrarily large table out as CSV, which is a great format for this because it can be generated without buffering the entire thing in memory. I wrote a bit about that here: https://simonwillison.net/2021/Jun/25/streaming-large-api-responses/ https://simonwillison.net/2021/Jun/25/streaming-large-api-re... zipfly makes me think that maybe I could do things like "stream all of the tables from this database as a zip file full of CSVs" - and have it work for giant databases again without using a great deal of memory. UPDATE: Actually it looks like zipfile wouldn't work here, because the library is designed to work from files that are already on disk - it doesn't look like I could feed it a few lines of CSV data at a time, I'd have to write the entire CSV out to disk first.
- password4321 5y agoSounds like you need this... a previous submission as linked elsewhere mentioned https://github.com/longaccess/python-zipstream/tree/streaminput https://github.com/longaccess/python-zipstream/tree/streamin... which should do it. Edit: Dang it, GPLv3, sorry (vs Apache 2 of your project). The two previous discussions with comments may offer enough to get you what you need: https://news.ycombinator.com/item?id=23198233 https://news.ycombinator.com/item?id=23198233 https://news.ycombinator.com/item?id=25616513 https://news.ycombinator.com/item?id=25616513 Edit 2: https://github.com/kbbdy/zipstream https://github.com/kbbdy/zipstream, no license (yet)
- d136o 5y agoInteresting, for the read (decompression) case I wrote this a while back: https://github.com/d136o/StreamingUnzip https://github.com/d136o/StreamingUnzip Basically, if you have a big zip file with many files in it (csvs for example), you can pipe out the decompressed data… It’s a bit obtuse to use since it calls for the end chunk of a zip archive (it may come from s3 for example).
- esjeon 5y agoI don't get why it has to be ZIP. AFAIK we can already stream tar-gz using only the python standard library. Maybe because Windows still can't unpack TAR out-of-box, 20 years after entering the 21 century?
- michalc 5y agoLooks good! I've been thinking about making a writable version of https://github.com/uktrade/stream-unzip https://github.com/uktrade/stream-unzip, but looks like you beat me to it! (Full disclosure: I'm the main developer of stream-unzip)