4 ms·
The idea that "if only web browsers didn't render bad HTML, the web would be so much better" is one of the oldest myths in web development, and I'm kind of amaz
by yoz 7y ago
The idea that "if only web browsers didn't render bad HTML, the web would be so much better" is one of the oldest myths in web development, and I'm kind of amazed that it's still showing up.
What you call the permissiveness of web browsers - in other words, their insistence on attempting to render invalid or badly-formed HTML - is what has made the web succeed at all.
Firstly, it was fundamental from the start: NCSA Mosaic was implemented that way, as was Netscape, so there's no point blaming Microsoft.
Secondly, and far more importantly, the robustness of web browsers is the reason why you can read 99.9% of web pages at all, including the one you're reading right now. (Yes, it's invalid: https://validator.w3.org/nu/?doc=https%3A%2F%2Fnews.ycombinator.com%2Fitem%3Fid%3D22146629 https://validator.w3.org/nu/?doc=https%3A%2F%2Fnews.ycombina... )
I know it's tempting to believe that draconian error handling would have forced people to code web pages "properly". Unfortunately, when draconian error handling was added to the web (as XHTML), it failed to take off. Check the history: https://www.w3.org/html/wg/wiki/DraconianErrorHandling https://www.w3.org/html/wg/wiki/DraconianErrorHandling
Mark Pilgrim wrote several excellent pieces about why non-draconian error handling is better, and as someone who wrote XML feed parsers and validators that were among the most robust and thorough in existence, he is deeply qualified to know. My favourite of those pieces is the "Thought Experiment"[1] but I also recommend [2], which includes:
There are no exceptions to Postel’s Law. Anyone who tries to tell you differently is probably a client-side developer who wants the entire world to change so that their life might be 0.00001% easier. The world doesn’t work that way.
[1] http://web.archive.org/web/20080609005748/http://diveintomark.org/archives/2004/01/14/thought_experiment http://web.archive.org/web/20080609005748/http://diveintomar...
[2] http://web.archive.org/web/20090306160434/http://diveintomark.org/archives/2004/01/08/postels-law http://web.archive.org/web/20090306160434/http://diveintomar...
- owl57 7y agoBut your publishing tool had a bug, and it automatically inserted their illegal characters into your carefully and validly authored page, and now all hell has broken loose. I'm not entirely sure an XML-error-deface is the worst way to expose a program that automatically inserts anyone's garbage in your web page while not having a clear model of acceptable garbage.
- s_gourichon 7y ago@yoz is right. If only web browsers didn't render bad HTML, the web would not be so much better, it would not have worked. There is a historical precedent to show it. I remember when XHTML was the future, about 2002-2005. Pages were loaded in Firefox by a XML parser. If the page was invalid XML for any reason, Firefox would render a parser error message: "error X in line Y, column Z" with a copy of the offending line and a nice caret under the error position thanks to a monospace font. Wrong percent encoding? No page rendered. Invalid entity? No page rendered. Messy comment separator (two minus signs)? No page rendered. Inserting an element where not allowed ? I guess no page rendered. This is nice for a rigorous developer perspective, I appreciated it. But (I used to hate that but a wise person sees the world as it is) it is a catastrophe for real-world adoption. Fixing one static page on your dev machine, thanks to the error message, is a thing. Making a dynamic website become practically impossible unless all your engineers are extremely rigorous and well-organized, and/or use a framework that generates guaranteed valid XHTML any time. But all frameworks (except a few unknown ones) had (have?) no notion of a document tree or proper escaping but just concatenate text snippets. From a business perspective, it means your website is much more difficult to get displayed at all (let alone correctly displayed). And even if it works today, it can blow up at any time because of a minor fix anywhere. Worse, the pages your team tests are okay, but real-world visitors will hit some corner case and get an error message intended for a developer. One may have hoped that some cleaner framework would appear and serve guaranteed valid XHTML any time. I would have liked this option. Developer would create tree hierarchies in memory and serialize them into XHTML. Please commenter name some that do and how popular they are. Did it save XHTML? That aspect may be the reason number one why XHTML was ditched in favor of HTML5: the web worked because it did a best effort to render invalid pages. Any solution that strays away from this principle will not be adopted at large. Meta bonus: we're discussing the HTML level but this kind of discussion we would have had at any other level, had the stack been consistent a few levels higher (script) or lower (HTTP, TCP). It's funny how HTTP and TCP looks like they just work, but they have their own corner cases and spec holes. The ecosystem just happened to have mostly converged on a few implementations that mostly work okay. (No, let's not talk about IPv4, NAT, and the like. ;-)
- zerotolerance 7y ago