4 ms·
I never understood that either except, as you said, for the hobbyist's sake. A lot of XHTML was generated from XML and, if one is using XML, chances are program
by sureaboutthis 7y ago
I never understood that either except, as you said, for the hobbyist's sake. A lot of XHTML was generated from XML and, if one is using XML, chances are programming is involved in the transformation to XHTML. But programming has strict rules itself and will also fail if not adhered to so I never understood the complaint of "draconian" error checking in XML/XHTML.
- gsnedders 7y ago> But programming has strict rules itself and will also fail if not adhered to so I never understood the complaint of "draconian" error checking in XML/XHTML. In most cases, if your code has some syntax error, the author of the code sees the syntax error; in the web case, if your code has some syntax error, the user sees the syntax error. That's the dramatic difference. The other reality is unlike program code, there's vastly more often user content intermixed in (X)HTML and it's rare for people to implement sanitisation correctly (do you handle U+0000? U+FFFF? U+1FFFF? most people outputting XHTML historically haven't, even if they get the security critical stuff (like "<") right).
- sureaboutthis 7y agoNot true. XHTML and HTML have validators to check your markup for proper syntax and usage. You know that. You made one of them! Good writers of markup will always check with those before they ship it. In the case of user supplied markup, that's still an issue today with HTML.
- gsnedders 7y agoI've been involved with multiple HTML and XML parsers, but never validators. :) The reality ten years ago, when a number of prominent XML advocates were using XHTML (and actually using it as such, serving it as such), almost all of their sites had user input means where the input was sanitized well enough for HTML to be secure (and not have any markup injection), but not for XML well-formedness (they got all the markup injection risks in XML, but not all the other WF requirements). If the very people who claim XML is easy can't get it right, can everyone else?
- sureaboutthis 7y agoYes. It was your outliner I was thinking of, not a validator but weren't you involved with the original http://validator.w3.org/nu/ http://validator.w3.org/nu/ at least in part? I, too, served my web pages as "real" xhtml 10 years ago and loved it :)
- gsnedders 7y agoPromise I've never worked on a validator!
- sureaboutthis 7y agoWell, now I have to rewrite my book :)
- jerf 7y agoAs the scale of a program increases, the probability that someone will do something wrong increases polynomially. Consequently, as web sites got larger and larger, the probability that some component would break the XML goes up quickly. This is a difficult pattern to deal with, pushing up the skill floor required, and as HTML5 shows, it isn't even all that necessary. There was also similar exposure from the data side; as the amount of data you handled increased, the odds that some data would tickle some code path that you didn't even know could blow up went up. You write your news front page in XHTML, and everything seems fine for six months, until someone finally includes an ampersand in their headline, and your entire front page crashes for two hours (not in any way monitoring will pick up, either, so you're getting customer reports), and it takes you hours to discover that someone was passing through the headline (and just the headline!) unencoded. The problem isn't XHTML's rigidity per se; personally I'm inclined more in that direction myself. The problem is when you have a ton of sloppy systems working together (MySQL, old HTML generation code, plugins from third parties your don't control, open source written by people whose belief in their understanding of HTML exceeds their actual understanding, decade-old internal databases with poor validations and unknown provenance, and so on and so on indefinitely), and then trying to suddenly, at the last minute, couple that big sloppy pile of technology to a strict technology at the last second. That sudden mismatch there at the end was a huge problem. One of the reasons I tend to prefer being as strict as possible is that in general, starting with existing strict-tech and adding sloppy-tech to it is no big deal; the sloppy tech doesn't complain that it only gets a subset of possible data it will accept. And if you need to couple to strict-tech, you still can. But if you start with a sloppy-tech system and for some reason need to couple it to a strict-tech system... prepare for some long nights and blown deadlines. So, professionally, the correct default is to choose strict-tech whenever possible. But XHTML forced that at almost the worst possible place.
- nostrademons 7y agoPeople underestimated the extent to which markup may be mixed in from sources you didn't control, and the power that this ability gives to your users. Say you built a shiny new forum engine with from-scratch XHTML markup. You try to sell it. Most of your customers say that they've already been running forum software since 1995 but the existing posts allow inline HTML (which was less unsafe in 1995, because no Javascript), which is all badly misnested. As soon as they dump the previous data, their site stops working. Or you import a small Javascript library from 1999 that generates its own innerHTML for a few elements, but does it with HTML. Oops. Or you built a new CMS with shiny XHTML markup, but before you had the CMS your org just hand-wrote pages which you now need to parse and import into the CMS. These were all very real considerations in the 2002-2004 period; I've dealt with all of them. Backwards compatibility is often the most important feature you can offer, because it directly affects the value the end user gets out of the product. Sites that were concerned with "doing it right" in that time period largely failed, while sites that "did it fast" in a hacky, XSS-prone way are now worth hundreds of billions of dollars.