3 ms·
I haven't read the rest of the thread, but the article is even wronger than the glib replies there. HTML needn't be well formed. There are adhoc rules which let
by 03b17999-4268 5y ago
I haven't read the rest of the thread, but the article is even wronger than the glib replies there. HTML needn't be well formed. There are adhoc rules which let major browsers parse broken HTML. If you do not follow the spec to the letter you will have your tooling break on input that every browser thinks is acceptable.
Which is why you always use whatever html parsing library comes with your language. There is no simple answer in the thread because there is no simple answer in the real world.
That said, anyone who says:
>It is quite possible and not even that difficult:
(
# match all tags in XHTML but capture only opening tags
<!-- .*? --> # comment
| <!\[CDATA\[ .*? \]\]> # CData section
| <!DOCTYPE ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* >
| <\? .*? \?> # xml declaration or processing instruction
| < \w+ ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* /> # self-closing tag
| < (?<tag> \w+ ) ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* > # opening tag - captured
| </ \w+ \s* > # end tag
)
Should be seriously mentored by someone.
- dgrunwald 5y agoBut both the original question and that article were about XHTML. The non-well-formed mess only matters for HTML, not XHTML. The regex is a valid answer to the stack overflow question.
- WesolyKubeczek 5y agoSorry, but no. The original question only mentions the author wants to ignore XHTML-style self-closing tags which is in no way implying the input to be well-formed XHTML.
- megous 5y ago> Should be seriously mentored by someone. Maybe people who think a basic regex such as this is difficult, need some mentoring. It doesn't even use lookahead/lookbehind or other more complicated features. It only uses non-greedy matching 2 times, which is the only thing that's more complicated than the basic AND/OR logic expressed by the rest of the expression.
- tannhaeuser 5y agoYou make this more complex than it actually is. HTML is basically SGML with a DTD declaring rules for tag omission/inference, empty elements, and attribute shortforms. This is what the majority of XML heads seem to struggle with, but which nevertheless is every bit as formal as XML is, XML being just a proper subset of SGML, though too small for parsing HTML. Then HTML has a number of historical quirks including: - the style and script elements have special rules for comments that would make older pre-CSS, pre-Javascript browsers just see markup comments and ignore those - the URL syntax using & ampersand characters needs to be treated special because & starts an entity reference in SGML's default concrete syntax - the HTML 5 spec added a procedural parsing algorithm (in addition to specifying a grammar-based spec) that basically parses every byte sequence as HTML via fallback rules; for most intents and purposes, the language recognized by these rules, taken to the extreme, is not what's commonly understood as HTML - WHATWG have added a number of element content rules on top of the HTML 4.01/HTML 5.0 baseline with ill-defined parsing rules (such as the content models for tables and description lists); the reason is precisely that WHATWG, once Ian Hickson had distilled the HTML DTD 4.01 grammar rules into the HTML 5 grammar presentation as prose, a formal basis was no longer used for vocabulary extension
- Rapzid 5y agoThe CDATA section doesn't appear to match this first example from Wikipedia: <![CDATA[<sender>John Smith</sender>]]> So it's either bugged due to the spaces or there is something else going on I don't understand; complexity. Agree on the sentiment. Through a "Simple made easy" lens, easy or hard, there is a lot of complexity in the task and this solution..
- jameshart 5y agoYou need to tell your regex engine to ignore whitespace in the expression. The regex handles that fine.