5 ms·
I built a streaming zip app using nothing more then the Python stdlib zip implementation and some os primitives. It runs on a small embedded device that can st
by 2bluesc 6y ago
I built a streaming zip app using nothing more then the Python stdlib zip implementation and some os primitives.
It runs on a small embedded device that can stream zip archives many times larger then the disk or system ram without any issue.
Example Python Falcon Proof of Concept:
https://gist.github.com/kylemanna/1e22bbf31b7e5ae84bbdfa32c68e03a9 https://gist.github.com/kylemanna/1e22bbf31b7e5ae84bbdfa32c6...
Other then what Python's zipfile buffers in memory, my implementation shouldn't use much more then a os.pipe()'s buffer (typically 64kB?).
- ejwhite 6y agoInteresting. I need to open a very large CSV file in Python, which is around 25GB in .zip format. Any idea how to do this in a streaming way, i.e. stopping after reading the first few thousand rows?
- 2bluesc 6y ago> I need to open a very large CSV file in Python, which is around 25GB in .zip format. Any idea how to do this in a streaming way, i.e. stopping after reading the first few thousand rows? Replace the `file_paths` list in my proof of concept with your large file(s), delete the rest (lines 61-68, 77-79) and it should just work.
- johndough 6y agoWorks fine with Python's standard library. Files in a ZipFile can be read in a streaming manner. There is no need to store all the data in memory. import io, csv, zipfile max_lines = 10 with zipfile.ZipFile("data.zip") as z: for info in z.infolist(): with z.open(info.filename) as f: reader = csv.reader(io.TextIOWrapper(f)) for i_line, line in enumerate(reader): if i_line >= max_lines: break print(line)
- 2bluesc 6y agoThis is true when writing to a file. The goal of my PoC was to not write a file and instead to stream to the web browser.