3 ms·
I almost wrote a zip parser. After studying the format a bit, I gave up that idea. - A valid zip does not begin with, as you would normally expect, a magic num
by twr 10y ago
I almost wrote a zip parser. After studying the format a bit, I gave up that idea.
- A valid zip does not begin with, as you would normally expect, a magic number.
- A valid zip can contain arbitrary prepended data.
- A valid zip can contain sections with no identifier value.
- A valid zip can contain arbitrary data between sections.
- A valid zip is validated starting with a tail section located at the end.
- A valid zip can have a valid tail section that contains a valid tail section.
- A valid zip can contain arbitrary appended data.
I don't get it: Why design a format that's so hard to parse? Implementing a single-pass streaming parser is impossible. It should be a basic requirement for most file formats. /usr/bin/unzip cannot even extract from standard input. I'm sure the implementer didn't feel like receiving user complaints about exhausted memory.
- badsectoracula 10y agoBecause the format was originally written by PKWARE for their PKZIP and PKUNZIP programs for DOS in the 80s and used for backup purposes, among others (being able to span multiple floppy disks was an often used feature). To write the directory at the beginning they'd need to keep all the data in memory or do some extra postprocessing that would make it prohibitively slow on a 4.7MHz machine with a 20MD hard disk and a slow 360K floppy disk drive. So the directory goes at the end. The files do not begin with a magic number because ZIP files can be embedded in other files - most commonly executable files for the self-extracting feature of PKZIP that placed the ZIP file right after the decompressor. This allowed the executable to be used both by itself and by PKZIP. As for the data after the magic number, it was used for the comment which was most likely added at a later point and they decided to put it after the magic number so that it remains compatible with existing decompressors. In the 80s and early 90s people didn't had Internet to get the latest version and a lot of old versions of PKUNZIP were floating around for many years.
- twr 10y agoThat is an excellent explanation, thank you.
- dkonofalski 10y agoTo further add on to that, I would frequently run into programs (and sometimes games) that would span multiple floppies where almost the entire first part of the archive was just installer data. Sometimes, the actual .zip file data would start on the tail end of the first disk and then span however many disks were needed while other times the first disk was just the installer and then the .zip data was on the other disks. In either case, PKZIP and PKUNZIP would be able to read that file but only in the 2nd case would they be able to read the file without that first disk. Things got further complicated when RARs made the scene because I remember that some of the early .RAR apps offered their own implementations of .zip support that would create multi-span .zips where the installer data was on the first disk, the .zip data was on the other disks, but the container spanned all the disks. This meant that you could potentially have 3 different ways of dealing with the same data and there was no safe way to assume which method was used: 1. Installer on disk 1, multi-spanned .zip on disks 2-n 2. Multi-span .zip on disks 1-n 3. Multi-span .zip on disks 1-n but installer data only on disk 1
- pixelglow 10y agoMark Adler, he of zlib/gzip/Info-Zip fame, seems to think that zips cannot contain arbitrary data before and between individual files. http://stackoverflow.com/a/12393597/60910 http://stackoverflow.com/a/12393597/60910 Therefore the straightforward way to parse a zip file is to proceed from the beginning and parse out each file sequentially. The End of Central Directory record is then only a redundant convenience to avoid sequentially scanning files e.g. in large zips for random access.
- twr 10y agoInteresting. It makes sense to be strict when parsing stream input. Skipping the redundant central directory section hadn't even occurred to me. Bonus points for eliminating the stupid comment confusion dilemma! On second thought, not entirely redundant, as the central directory does contain the file permissions. But those can be parsed and set after file extraction, without increasing the overall memory complexity.