8 ms·
This strikes me of an example of a sort of dualism that exists within software. The .zip spec is flawed, and can be broken in all sorts of different ways. Howev
by jplasmeier 10y ago
This strikes me of an example of a sort of dualism that exists within software. The .zip spec is flawed, and can be broken in all sorts of different ways. However, after years and years of using computers, I've never had an issue with extracting or using a .zip file.
I wonder if this same pattern can be applied to physical engineering disciplines- e.g. a structural engineer assessing a bridge and finding numerous faults in the design, despite traffic still using the bridge as normal...
- niftich 10y agoMost out-in-the-wild usage of zip files fits the 'Robustness Principle', or, after its author, Postel's Law, found in RFC 1122 [1]: > Be liberal in what you accept, and conservative in what you send There are several analyses of this maxim and whether it's a good choice for designing robust, secure systems. This particular Internet Draft doesn't agree [2]. [1] https://tools.ietf.org/html/rfc1122#section-1.2.2 https://tools.ietf.org/html/rfc1122#section-1.2.2 [2] https://tools.ietf.org/html/draft-thomson-postel-was-wrong-00 https://tools.ietf.org/html/draft-thomson-postel-was-wrong-0...
- user5994461 10y agoFor user interfaces: Be liberal in what you accept from the user. For M2M (machine to machine): Be strict in what you accept, and strict in what you send. That's my take on it now. For instance, a user will copy/paste URLs to his browser, there will probably be a space too much before or after. It's okay to clean it (and a better experience than a "site not found"). When a machine sends something "weird". Well, it's not possible to know if it's really wrong and it can't be corrected in any meaningful way. Just fail and throw an error so the developer can fix it.
- wolrah 10y agoI agree with those positions, but what about the middle ground? Browsers for example, taking machine-to-machine data which sometimes is human generated. I'm a fan of the strict solution as far as that goes, but there's a reason XHTML failed. Somehow asking people who write web sites to do it right or else it doesn't work is a big deal.
- taneq 10y agoThe best approach seems to be to be strict in what you accept, and strict in what you emit, and if you want your software to be forgiving to the vagaries of human input, then strictly adhere to a forgiving spec.
- Too 10y agoThe difference is a user interface translates the non strict version to the strict version instantly once you click the submit button. A hand written html file is stored persistently, retaining the flaw for all future.
- user5994461 10y agoWell. How to say that without being dismissive of the average web hobbyist... Well, when the job is to accept a hugely complex flexible poorly defined input format from 30 years ago that is written by millions of people who have no clue what they are doing, fixing common errors and getting anything to render is part of the spec ^^ And... oh wait! I said "Be liberal in what you accept from the user" when it comes to user input. Web pages are user input, so yeah, browsers are no exception to the rules, they're actually a perfect example! :D
- chrismorgan 10y agoIf the web had been strict from the start, we would never have had a problem: people would never have been publishing invalid documents because browsers would have been rejecting them outright. XHTML failed because HTML existed and was good enough.
- taeric 10y agoNo. You should always be as liberal on what you accept as you can. Primarily because as strict as you try to be in what you send, you are likely to make mistakes. Very very likely. Obviously, your field of work dictates a lot here. And, if you are accepting something that has severe consequences on acting out, then yes, be more strict. However, the general principal holds. In general.
- user5994461 10y ago> Primarily because as strict as you try to be in what you send, you are likely to make mistakes. Very very likely. What you call a human mistake should be called a bug. The only acceptable way to handle a bug is to fix the code. Not to be dismissive but the way you think is a classic beginner mistake. It's not the responsibility of other software to guess what bugs you'll put in yours. ^^ > Obviously, your field of work dictates a lot here. large distributed systems, financial exchanges, trading systems, national government projects, aerospace, and even web stuff at times. Some fields have low standards. That doesn't mean that the good practises don't apply, it just mean that people don't apply them ;)
- taeric 10y agoI conceded it is a bug. I merely claimed you will make some. Because, you will. Period. You also have an odd misconception. It isn't having lower standards to be liberal in what you accept. It is actually harder. Much harder. And you should do it. Easy example from finance. You shouldn't just take one currency. You should take in as many currencies as you can understand. Does this mean to be magical? No. But you should ask, "how many different ways could this be given to me?" And you should instrument this with a marker for "unexpected."
- user5994461 10y ago> You shouldn't just take one currency. You should take in as many currencies as you can understand. This is a business feature request. The "Being liberal about what you accept" is a technical guideline for protocol/format design and input processing/sanitation. Don't apply that rule to feature requests, it's not meant for that :(
- vacri 10y agoAs an aside, a couple of months ago I found out that the PHP ISO8601 timestamp format is not 8601 compliant. If you want an ISO8601 timestamp, you have to use a different format. I'm not sure how this fits in with Postel's law, but your comment jogged my memory of it :) http://php.net/manual/en/class.datetime.php http://php.net/manual/en/class.datetime.php
- Terr_ 10y agoMySQL utf8 isn't utf8 either :p
- flukus 10y agoI wonder how many security exploits this principle is responsible for?
- manarth 10y agoThe .zip spec is flawed, but… I've never had an issue with extracting or using a .zip file I'd suggest this is thanks to most software following (to some extent) the robustness principal: be conservative in what you do, be liberal in what you accept. Most of us will typically encounter fairly well packaged, conforming zip files. Occasionally we may come across something unanticipated - like this example, where HTML content is accidentally appended to the end of a zip file - and I suspect this is where we will find ambiguous behaviour: some package tools may crash, others might "extract" it as though it were content, others might ignore it. It's this area of ambiguity that lends itself to vulnerabilities and attacks. Re: physical engineering, I'd recommend a great book: "To Engineer Is Human". It talks about the evolution of engineering, which is a surprising amount of trial and error, with emphasis on the error.
- mohaine 10y agoSorta like how your keys are always in the last place you look. Everything before good enough is an error. Good enough is good enough. This is engineering.
- TAForObvReasons 10y agoCSV is a great example of this phenomenon. There is a "spec" RFC4180 and there are tools that generate CSV files that do not technically conform to the spec. One such tool is Excel. For most users, Excel is doing the right thing. Blaming Excel for not handling CSV files according to the spec is passing the buck. The CSV tools that are worth using generally take great pains to work with Excel files at the cost of some ideological purity. IMHO it's a reflection of the software developers involved. The best tools, the ones we turn to time and time again, generally just work.
- niftich 10y agoIn fairness, usage of CSV-like formats pre-dates the CSV RFC by almost 30 years, which was authored in 2005 specifically to try to formalize a de-facto spec: > Surprisingly, while this format is very common, it has never been formally documented. [1] [1] https://tools.ietf.org/html/rfc4180#section-1 https://tools.ietf.org/html/rfc4180#section-1
- fnord123 10y ago> For most users, Excel is doing the right thing. Blaming Excel for not handling CSV files according to the spec is passing the buck. No. Excel is wrong when it comes to CSV. Paste a Unicode string into Excel. e.g. Beijing in Simplified Chinese (北京市). Now Save As Windows CSV as beijing.csv. Close the file. Open beijing.csv. The cell now reads `___` (on Excel for Mac 2011 - maybe they deigned to fix it). Excel just outputs bad data.
- jefffan241 10y agoI don't know how you do it in excel but if you generate a csv with UT8 data you can add a byte order mark[0] as the first byte and it will render correctly. Once you add that, excel will open the file with utf8 encoding (if you use the utf8 byte order mark obviously). I haven't tried with other utf-* encodings. Again don't know how to tell excel how to add that though :/ I've only had to deal with arabic in generated csv's. [0] https://en.wikipedia.org/wiki/Byte_order_mark https://en.wikipedia.org/wiki/Byte_order_mark
- amelius 10y agoWell, you have to keep in mind that mistakes like this could turn into exploits.
- teaearlgraycold 10y agoDidn't expect to find you on hackernews, plaz
- asmithmd1 10y agoYes, this happened with a building built in NYC in the late 1970's: http://www.slate.com/blogs/the_eye/2014/04/17/the_citicorp_tower_design_flaw_that_could_have_wiped_out_the_skyscraper.html http://www.slate.com/blogs/the_eye/2014/04/17/the_citicorp_t... An undergraduate doing a class project on the building uncovered the flaw.
- mmahemoff 10y ago"It was only after seeing the documentary that she began to learn about the impact that her undergraduate thesis had on the fate of Manhattan." Amazing story.
- xg15 10y agoThis feels kind of like concluding that the y2k bug was overblown because in the end nothing happened - and ignoring the reasons why nothing happened. Those dualisms can usually be resolved if you realize the vast and complex efforts that go into working around all the spec bugs - in this case, the various "find the magic number" heuristics.
- optionalparens 10y agoI have indeed encountered many broken zip files. In truth it doesn't happen to me much anymore, bit it sure did in the earlier days of pkunzip/pkzip, and in general during the BBS days. Sometimes this was just due to all sorts of weird things going on with transmitting files over networks and modems. Mostly you could prevent these things with CRC checks or any other verification method, but that doesn't mean implementations checked this or that the check itself didn't just fail. A few things that I've hit in the wild that screwed up zips: * Viruses/Worms that were rather primitive and start adding weird bytes in all kinds of places in all kinds of files * Incomplete transmission over a modem. Depending on the protocol, you might even have most of the data, so you could read part of zip header, but not the actual archive or vice-versa. Normally, you knew right away the file was incomplete with a CRC check so not the worst problem, but the check itself was slow on old hardware. * Weird things that added metadata and screwed with byte order, ends of file marks, etc. For instance there were some early attempts in the BBS scene to add metadata formats similar to what ID3 is to mp3s. Sometimes the writers would ruin the original file. I hit a few cases of strange attempts at steganography with software pirates trying to be "3l337" or whatever. * Tools people tried to use to fix broken zips, that didn't quite fix them how they thought. * Not really the fault of zip, but I've seen people rename a zip's extension to another archive format, causing the unarchiver to assume the format based on the extension. Never trust an extension if you can help it. (ex: rename a zip to rar, arc, lzz/lha, tar whatever) * Floppy and hard disk repair programs. Sometimes these things would end up corrupting the bytes of files instead of fixing them when they tried to be more clever than moving things around. Sometimes moving things around also would result in things being ordered wrong for whatever reason in these programs. Some of the DOS Norton/Fastback/etc. ilk were especially frequent offenders.
- taneq 10y agoI believe this is an example of the old adage, "In theory, there's no difference between theory and practice. In practice, however, there is." In theory, the .zip spec is broken. In practice, it's the most reliable format for transferring a group of files. (And don't even get me started on JPEG, where iirc the file format wasn't even specified until after JPEG files had been popular for years.) > I wonder if this same pattern can be applied to physical engineering disciplines- e.g. a structural engineer assessing a bridge and finding numerous faults in the design, despite traffic still using the bridge as normal... I would be amazed if this weren't the case. I know I've encountered a few cases where mechanical engineers had cocked up the design but the resulting machine still managed to limp along and mostly perform its function.
- netheril96 10y ago> In practice, it's the most reliable format for transferring a group of files. For those who speak only English, yes. The encoding of filenames have always been a mess for zip files.
- taneq 10y agoTrue, that's a bit of a blind spot for me, being an only-English-speaker.
- user5994461 10y ago> I wonder if this same pattern can be applied to physical engineering disciplines Oh yes. Basically, EVERYTHING is flawed in INFINITE ways. Or to put it differently, perfection doesn't exist. Luckily for us, the barrier for "being practical and useful" is a lot lower than perfection =)