12 ms·
HTML Optional Tags
- paulddraper 3y agoAnd people wonder why XHTML was a thing. (Is still a thing, but W3C recommends against it)
- JodieBenitez 3y agoYes... but why do this ? I don't regret the XHTML days and its feature stagnation, but this is just useless.
- lifthrasiir 3y agoIt is a formal specification of what browsers used to do with a broken HTML. Because it is now fully specified, everyone can safely write a broken HTML (half joking!), but also there is no longer a surprising behavior due to diverging behaviors from different browsers.
- JodieBenitez 3y agoGood... now let's make a formal specification for this: <ahahah>mmm... <ohoho>nope.</ahaha></ohoho> Fully joking :-P
- ZeroGravitas 3y agoI believe whatwg did in fact specify for this kind of overlapping tag. edit: apparently it has the cute name of "the adoption agency algorithm": https://html.spec.whatwg.org/multipage/parsing.html#adoption-agency-algorithm https://html.spec.whatwg.org/multipage/parsing.html#adoption... > Note: This algorithm's name, the "adoption agency algorithm", comes from the way it causes elements to change parents, and is in contrast with other possible algorithms for dealing with misnested content.
- ainar-g 3y agoAt least the people who named it understood the silliness of the issue, heh.
- lifthrasiir 3y agoIt is not as silly as it sounds. Many programming language implementations allow a partial parser that looks similar to this in order to give better diagnostics even when the AST is incomplete. (I did the same in the past, for example.)
- MrVandemar 3y agoIt's very useful. I use optional -- therefore abbreviated -- HTML syntax as an alternative to MarkDown for writing.
- JodieBenitez 3y agoYou're saying you're using an abbreviated HTML as an alternative to a markup language that is a lightweigth alternative to the markup language known as HTML ? We need to go deeper. •`_´•
- MrVandemar 3y agoYes, exactly. There are three advantages: 1. No translation of markdown to HTML is required to turn it into a web-page, the medium I primarily work in. It's not required -- I can read nicely formatted raw HTML just fine. 2. Lower cognitive load. Damn if I can remember if a link in markdown is ()[] or []() or the link goes first or the link goes last. But a link in HTML is consistent with the rest of the language. 3. HTML is richer than Markdown. I make heavy use of <abbr> and <cite> and <time>, not to mention metadata tags. My position is if you need to include HTML in Markdown to express yourself, you might as well write in HTML instead of construct some hybrid bastard document.
- JodieBenitez 3y ago> My position is if you need to include HTML in Markdown to express yourself, you might as well write in HTML instead of construct some hybrid bastard document. I agree. I just don't think it's very useful to omit close tags. Any decent editor will maintain them for you.
- GoblinSlayer 3y agoThis way you can send spam sms messages with little html markup and google analytics.
- jraph 3y agoI learned HTML when XHTML 1.0 was current. I've long preferred strict HTML written using XML. It helps spotting mistakes, and there are no parsing surprise. Now that HTML5 parsing is well specified, I've come to think that either you want to be strict and have the browser tell you something is wrong, and you use XHTML for this, or all these optional tags are just useless. I want to optimize readability, and then file size. I believe closing all tags you opened and quoting all attributes helps readability, and also that all these <head>, <body>, <html> tags just get in the way and make your eyes go through useless boilerplate and makes your fingers type useless things too if you don't use templates. You still need to specify the charset so characters are interpreted correctly, so for me, if you are not going to use application/html+xml anyway this works well: <!DOCTYPE html> <meta charset="utf-8" /> <title> My title </title> <p> Lorem ipsum... </p> Both quicker to read and write, while not raising maintenance costs. Though just yesterday I edited my resume written in XHTML and the browser actually spotted a dumb mistake, so I still like the strictness of the XML parsing mode. One counterpoint to dropping the optional tags is for pedagogy: if I had to teach HTML to someone, I would make them use all the tags, or the result of having html and body in the DOM and CSS working on them will be very confusing. Only when they understand the DOM, what nodes are in an HTML page, I'd make them drop the tags if they want. Which is an important step so they can understand that nodes that are present in the DOM are not necessarily in the source code.
- layer8 3y agoIt was a codification of existing browser practice. Specifying something different wouldn’t have changed browser behavior, it would only have led to browsers ignoring the spec.
- pwdisswordfishc 3y agoHTML is still feature-stagnated; features are mostly added to JavaScript and DOM APIs.
- JodieBenitez 3y agoYeah, let's say I'm late to the party then. But still, the improvements compared to HTML4 are huge.
- jimmaswell 3y agoCSS gets relatively frequent new features too. I think this status quo is fair enough - HTML is just there, and I never really think about needing more out of it. CSS and JS though I do find myself waiting for browsers to support upcoming experimental features often enough.
- alerighi 3y agoWhy not. If there is no ambiguity you save characeters (few bytes, but for each page) and thus pages will load faster even on slow connections. Also if you write HTML by hand (something not a lot of people does these days, but for example for my site I do) it's less characters to type and it's simpler.
- kevincox 3y agoThe problem is that every parser and emitter needs to be aware of these weird and changing rules. It wouldn't be that bad if the only things that read HTML were browsers but as it is every language has HTML parsers that are broken different ways leading to bugs and security vulnerabilities. For example emitters need to know what void elements are because <br></br> is actually equivalent to`<br><br>`. But `<script src=foo.js/>` is only an opening tag so the rest of your document will be executed as JavaScript. So you can't just write an emitter for arbitrary elements, you need to emit different things for `br` and `script`. Plus `script` has special escaping rules that are often forgotten about. Plus you better keep that list up to date! With XHTML it is very easy to write a parser that will construct a tree forever and can reserialize it with no issues. I have no issues with consistent changes such as empty attributes and unquoted attribute values, but I think that these element insertion, auto-closing, void elements and non-replaceable character data are a mistake because you need to maintain an up-to-date dataset of these custom rules or you get an incorrect result.
- throwaway87651 3y agoFantastic! These are great suggestions to help write readable, maintainable HTML. Similar to: https://lofi.limo/blog/write-html-right https://lofi.limo/blog/write-html-right
- danbruc 3y agoWould there have been a way to avoid this mess? - a browser must reject any invalid HTML in order to force the developers to fix their HTML - a browser must try hard to make sense of messed up HTML, otherwise users will switch to a competing browser that renders the mess for them Theoretically all browser vendors could coordinate so that everyone rejects invalid HTML, but there is probably no good way to avoid defectors. Why did this not happen for other technologies? My first thought was that there is no compilation step which allows forcing the developer to fix things without giving the end user any power through their choice of browser. But that seems not quite right, why do Bash or Python or your C++ compiler not make a best guess what your code is supposed to do? Because there is or was only one dominant implementation and therefore no competition? Because document markup is much more robust against small errors and probably remains readable while your code likely just crashes? That is probably one of the most important ones, I think. What role did browser specific features, evolving standards and incomplete implementations play? What is the end result? Nothing for the end user, they do not care whether the browser has to deal with nice HTML or a mess. Developer writing HTML get to be more sloppy at the price of a lot of additional complexity and pain where ever code has to deal with HTML. This might actually have some negative impact on end users because of bugs or security issues stemming from the additional complexity. Maybe it made HTML somewhat more accessible to the casual user as they could get away with some mistakes. But was this worth it, could better tooling not have achieved the same with good error messages helping to fix errors?
- lifthrasiir 3y agoIt is not really the mess---see the next chapter, 13.2 Parsing HTML documents, to see the actual mess. In fact, the HTML specification defines two concrete syntaxes for HTML where the first one is for `text/html` and another is for `application/xhtml+xml`. The latter has been never deprecated (thought the name XHTML was abolished). Moreover the spec states that: > Some authors find it helpful to be in the practice of always quoting all attributes and always including all optional tags, preferring the consistency derived from such custom over the minor benefits of terseness afforded by making use of the flexibility of the HTML syntax. To aid such authors, conformance checkers can provide modes of operation wherein such conventions are enforced. In the other words it recognizes the benefit from explicit tags, but also recognizes the benefit from optional tags. So they are equally conforming.
- hyperhello 3y agoThere is some variant of the theorem about any sufficiently complex language can’t express its own correctness and vice versa. We want to turn all expression failures into syntax errors; but we can’t. Just don’t write bad HTML, bad JavaScript, bad CSS, and there won’t be any trouble for you.
- mattkenefick 3y agoIt bummed me out when modern browsers started supporting mistakes in code; like when Chrome would interpret incorrect markup and fix it for you.