4 ms·
The problem with XHTML is that the "abort on parse failure" behaviour simplifies a computer science problem at the expense of creating a business problem; now i
by jgraham 6y ago
The problem with XHTML is that the "abort on parse failure" behaviour simplifies a computer science problem at the expense of creating a business problem; now if you have an error somewhere in your content generation pipeline it means that the site goes down. That's a pretty difficult tradeoff given that most CMS are absoutely not designed in a way that ensures well formed markup.
Back in the dim and distant past when XHTML was a fashionable term to throw around, Evan Goer did a study of whether sites claiming to serve XHTML were actually doing so. The results were not pretty [1]. Some people took the results as a challenge and tried to ensure they were sending valid XHTML with the correct mime type so that browsers would catch fire in the case of a parsing failure. In almost every case it turned out to be possible to get their sites to break with user generated content (e.g. searchng for XML-invalid characters which were then echoed back onto the page).
So I contend we ran the XHTML experiment pretty thoroughly 15 years ago, and it turns out that it doesn't really work. Once you accept that parsing errors being fatal isn't viable for publishing, you have to have some kind of error recovery system. The one in HTML isn't ideal since it's basically just the codification of many years of improvisation and reverse engineering. Maybe something like XML5 [2] would be better. But figuring out how to move the world in that direction is an unsolved problem. Meanwhile HTML Just Works for most of the people most of the time.
[1] https://goer.org/Journal/2003/04/the_xhtml_100.html https://goer.org/Journal/2003/04/the_xhtml_100.html
[2] https://github.com/annevk/xml5 https://github.com/annevk/xml5
- pwdisswordfish4 6y agoThe right solution is pervasive programming language support for interpolation inside XML/HTML code. Like there is in JavaScript now, known as JSX. Like HHVM has, in the form of XHP. (Interesting that both are Facebook innovations.) The industry didn't defeat SQL injection through more permissive query parsers, but by generating queries via ORMs and parametrised templates instead of dumb string concatenation. Also, have you noticed that JSON injection bugs are almost never heard of? That's because in many programming languages this very problem comes pretty much pre-solved before you even add any JSON support to them.
- sergeykish 6y agoError is still there, it is just a non breaking error. Imagine Word or Excel document that's trying to silently recover. I do not see programming languages adopting "not to fail" approach 1 + "2" //"12" 1 - "2" -1 There were PHP sites with mysql connection error all around. As industry we've chosen AirBrake approach — fail and notify developers. HTML makes it easier to edit plain text but there is a price. What you load is not what you've stored, HTML is a lossy serialization [1] [2] [3]. Program not human produced DOM, it should be safe to serialize-deserialize. It could be JSON, XML, s-expressions. It is unsafe with HTML. It is very easy to author XHTML. Start DOM first (HTML for brevity): data:text/html;charset=UTF-8,<p contenteditable>foo Done, it automatically escapes <>&. Extend with some controls [4], store it as XHTML [5]. It is WYSIWYG, much easier than HTML authoring. [1] http://sergeykish.com/script-style-is-cdata-in-html http://sergeykish.com/script-style-is-cdata-in-html [2] http://sergeykish.com/content-after-html-appended-to-body-in-html http://sergeykish.com/content-after-html-appended-to-body-in... [3] http://sergeykish.com/pre-newline-ignored-in-html-test http://sergeykish.com/pre-newline-ignored-in-html-test [4] http://sergeykish.com/live-pages http://sergeykish.com/live-pages [5] http://sergeykish.com/bookmarklet-put-xhtml http://sergeykish.com/bookmarklet-put-xhtml
- barrkel 6y agoHTML mostly ends up rendering text and images, and you have to mess up really hard to lose both of those. Typically what breaks is styling. It might not be pretty, but it may still be functional.
- sergeykish 6y agoI can't mess it when I edit DOM and browser restores it as it was. I may have <ul> in <p> (we had it in 1978), I may have <a> in <script> (and it works like comment), I may have <pre>\n and don't worry that it disappear each time I save document. I may have nested <script type="foo"> tags [1]. DOM supports it. XHTML supports it. HTML breaks my content on save-load. [1] https://stackoverflow.com/a/59548670/5554075 https://stackoverflow.com/a/59548670/5554075