11 ms·
Solving the regex of madness, and snarky answers on StackOverflow (2019)
- math-dev 5y agoGreat article (assuming the solution provided works). I do a lot of parsing in my projects, I find natural text based input vital for power users who don’t to point and click always. What are some good parsing algorithms, theoretical articles etc to help me become more professional in the parsing tools I write?
- deleted 5y ago[deleted]
- shmageggy 5y agoIt's not explicitly stated, but I believe the author's point is that the original question didn't require a recursive solution (because it's only asking about individual tags, not matching opening tags with their closing partners) Edit: yes looking at the answers, someone pointed this out in a comment response to the"Chomsky" answer: > The OP is asking to parse a very limited subset of XHTML: start tags. What makes (X)HTML a CFG is its potential to have elements between the start and end tags of other elements (as in a grammar rule A -> s A e). (X)HTML does not have this property within a start tag: a start tag cannot contain other start tags. The subset that the OP is trying to parse is not a CFG. – LarsH Mar 2 '12 at 8:43
- inopinatus 5y agoLook again. The context is definitely one of identifying tags that should be closed, but aren’t, and being identified thus in order to normalise a malformed document into well-formed XHTML. It is not prominently stated, but it is nevertheless the case. Consequently, balance matters. The clue is in the title: “match open tags”, for which the body text has only one part of the author’s train of thought (i.e identifying the opening tag). Further discussion of matching the closing tag is omitted from the body of the question, and many people forget (or disregard) the nuanced difference in the title at this point in their read-through, not stopping to ponder “why the heck would they just want the opening tags? what have I missed here about the context?” and instead treating it like some particularly awful and weirdly contrived exam question, as in the article at the top. The final confirmation of all this is in the question history of the question author around the same date: it’s definitely what they were working on. Consequently, our zalgo-spewing correspondent has it right, they should use a parser, and in addition the question should’ve received feedback early to help them describe the context more clearly. That all this was never properly clarified is a failure of moderation, and further a demonstration of how many folks struggle to challenge (or even identify) their own assumptions.
- goto11 5y ago"Match open tags" obviously refers to writing a regular expression which matches opening tags, not to pair opening tags with end tags. If you look at the regex which the OP themself suggests, it is clearly only intended to match opening tags (excluding self-closing tags), not search for corresponding end-tags.
- inopinatus 5y agoWell, as now expressed at hopefully sufficient length, it doesn’t just say solely that, unless one a) disregards the difference in phrasing, and then b) disengages any sense of purpose and practical utility and instead treats it like a badly worded test question. It’s kind of a shibboleth, in a way, for developer sensitivity to actual needs, as opposed to getting hung up on how clever they are.
- jcelerier 5y ago> disengages any sense of purpose and practical utility and instead treats it like a badly worded test question. Who are you to know the purpose and utility better than the person who asked the question ?
- inopinatus 5y agoOh, are they here? It’d be fascinating to hear from them. Alternatively, perhap, that hostile tone is suggesting I’m personally unqualified to interpret loosely framed questions? I suppose, since I’ve only been doing it for a few decades, I’m definitely a novice by any standard, and my tendency to observe and follow up on anomalous, incomplete, subtly conflicting, or otherwise inexplicable requirements by investigating both the timeline and substance of the original context, and the apparent motivations and outcome preferences of the author, is sheer beginners luck, and any uptick in stakeholder utility that from time to time accompanies amending recommendations following such investigative and analytical activity a sheer coincidence! So that must be it - as you can probably tell from all this smug, empty bravado, I’m really just sharing pure speculation, wild guesswork, total fluke, impertinent leaps of inferential faith, only just grasping at the vague outline of my own blind spots et cetera et cetera, and consequently yes, I’d love to hear the original intent restated from the horse’s mouth, too; but, for the meantime, I’ll read the tealeaves, systematically analyse and synthesise to the best of my ability, describe and discuss any substantive points of comprehension that I think might help enrich, or at any rate challenge, a reader’s perspective (including my own), and cross my fingers hoping to read, mark, and inwardly digest what new understanding or revealed wisdom as I can - even when it comes, as it has there and here, via diacritical allegory and dialectical hellfire. Or, finally, if you just want the TL;DR version, the reason I feel unshakeably comfortable asserting that the question author's actual purpose is normalising a nonconforming document into XHTML, by balancing open & close tags, is because they said so. It's in an answer comment about halfway down the first page. > "Can you provide a little more information on the problem you're trying to solve" "Determining all the tags that are currently open, then compare that against the closed tags in a separate array" If that's not enough, here are quotes from their other questions, posted in the minutes and hours prior: "How to balance tags with PHP", and "I need to write a function which closes the opened HTML tags." Sometimes, we just have to bother reading what's in front of us.
- 03b17999-4268 5y agoI haven't read the rest of the thread, but the article is even wronger than the glib replies there. HTML needn't be well formed. There are adhoc rules which let major browsers parse broken HTML. If you do not follow the spec to the letter you will have your tooling break on input that every browser thinks is acceptable. Which is why you always use whatever html parsing library comes with your language. There is no simple answer in the thread because there is no simple answer in the real world. That said, anyone who says: >It is quite possible and not even that difficult: ( # match all tags in XHTML but capture only opening tags <!-- .*? --> # comment | <!\[CDATA\[ .*? \]\]> # CData section | <!DOCTYPE ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* > | <\? .*? \?> # xml declaration or processing instruction | < \w+ ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* /> # self-closing tag | < (?<tag> \w+ ) ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* > # opening tag - captured | </ \w+ \s* > # end tag ) Should be seriously mentored by someone.
- dgrunwald 5y agoBut both the original question and that article were about XHTML. The non-well-formed mess only matters for HTML, not XHTML. The regex is a valid answer to the stack overflow question.
- WesolyKubeczek 5y agoSorry, but no. The original question only mentions the author wants to ignore XHTML-style self-closing tags which is in no way implying the input to be well-formed XHTML.
- megous 5y ago> Should be seriously mentored by someone. Maybe people who think a basic regex such as this is difficult, need some mentoring. It doesn't even use lookahead/lookbehind or other more complicated features. It only uses non-greedy matching 2 times, which is the only thing that's more complicated than the basic AND/OR logic expressed by the rest of the expression.
- tannhaeuser 5y ago
- dataflow 5y agoTry that regex on < script> console.log("<script2>"); </script> Edit 1: I'm unsure if the inner <script2> is valid (X)HTML, so it might not be an issue of being unable to parse correct (X)HTML, but rather an issue of being unable to detect invalid (X)HTML. (Can someone verify?) Edit 2: It seems Chrome chokes on the space... does anyone know if the initial space is valid? I'm pretty sure I've seen parsers that accept it...
- adament 5y agoBut it should be easy based on this example to include correct HTML tags in the script which the regular expression will emit. Or if you want to recognise HTML tags in the script, you can easily obfuscate construction of in the script using string concatenation.
- eCa 5y agoI don’t think any part of that is valid XML. There cant be space between < and the tag name, and I believe content containing tags should be in a CDATA section.
- goto11 5y agoThe question is about XHTML, not HTML. HTML does not even have the self-closing tags the question is concerned about. In XHTML, either the opening angle bracket must be escaped, or the script should be in a CDATA section.
- capableweb 5y ago> Edit 2: It seems Chrome chokes on the space... does anyone know if the initial space is valid? I'm pretty sure I've seen parsers that accept it... Most browsers probably can deal with it, but it's not valid xml/html. Try passing it through a validator, it'll complain about foreign characters after `<` and then complain about a trailing `</script>` as <script> was never opened in the first place.
- adament 5y agoDoes the proposed regular expression really handle embedded script content correctly? From my limited understanding of HTML, pretty much only </script> counts as closing the script contents and everything else is treated as part of the script.
- goto11 5y agoThe question is about XHTML though, not HTML which have a more complex syntax.
- jontro 5y agoTo me it’s a bit ambiguous if the original question is about both html and xhtml. It’s tagged with both
- goto11 5y agoThe headline says XHTML. HTML does not even have the self-closing tags the question is concerned about.
- jameshart 5y agoThe headline said that the author was specifically looking to avoid 'xhtml self-contained tags', but that doesn't mean they could assume the document they were looking in was valid XHTML - just that it might contain XHTML-style 'self-closing' tags. I've seen plenty of plain HTML files that aren't well-formed XML yet contain <br/> tags.
- goto11 5y agoYeah the trailing slash in <br /> is legal in HTML, but it doesn't actually make the tag self-closing. For example <b /> is still an opening tag which require a </b>.
- megous 5y agoDoes it matter? You can create regex with lookahead for </script>. The point is that it's possible to solve the problem from SO this way, due to the nature of the problem, not that this particular expression is perfectly correct.
- lifthrasiir 5y agoThe only part I agree in this writing is that you don't need to be snarky to be correct. (I'd like to introduce the XY problem of the second kind, where the answerer is so confident that it is the answerer who have missed the actual question.) Some regexes can recognize a language beyond the regular language. They are typically available in two flavors: recursive references (Perl, Ruby, PCRE) and stackable captures (.NET). They are obscure enough that I would not recommend them, but it is patently false that regular expressions (EDIT: of the practical interest) cannot be recursive. It is possible to match individual HTML tags with regexes, but it is difficult. It cannot use a bare `\w` or `\s` because both XML/XHTML and HTML5 parsers have peculiar definitions for tag name characters and space characters. For example your `\s` will typically match various Unicode space characters, while only ASCII whitespaces are recognized in tags. There are also several notable exceptions to the parser (and external states termed the "tree construction"), so missing any of them would result in an immediate XSS. If you think you can write a correct regex for HTML tags, my quizzes [1] should make you concerned. Limiting the question to XHTML does alleviate some but not all concerns. The distinction between recognition and parsing is correct, but parsing doesn't necessarily mean the reconstruction of parse tree. Parsing means the access to constituent nonterminals, which can be used to reconstruct parse tree but also directly used as their own (e.g. calculators). Indeed in most regex implementations you can't extract two or more strings out of each capture (Raku is a notable exception), so you can match against e.g. `(\w+)(?:,(\w+))*` but can't extract a list of comma-separated words with it. Practically speaking this means you can't extract a list of attributes with a single regex, making it unsuitable for HTML parsing anyway. [1] https://news.ycombinator.com/item?id=26355451 https://news.ycombinator.com/item?id=26355451
- dataflow 5y ago> it is patently false that regular expressions cannot be recursive. No, it's more like the term "regular expression" has gotten hijacked and nowadays gets abused to colloquially include, shall we say, irregular expressions. i.e. people basically say "regex" when they mean "some succinct pattern language with syntax similarities to (classical) regular expressions".
- 5y ago
- NtrllyIntrstd 5y agoThe author seems to be missing the point, in my opinion. While it is certainly true that often one can solve simple, seemingly innocent sub-problems within more general languages, the transitions from "I see I can solve this simple program with regex'es!" to "Then I can probably solve this other, almost identical problem as well!" and have the problem explode right into your face are subtle (almost imperceivable to a novice) and it would be a more robust solution to go for the right tools (i.e. an (x)html parser), as well as a good learning example. On a side note: regular expressions can not - by definition - parse recursive languages. A regular expression matcher that does is not a regular expression parser but an ugly-duckling in the family of context-free grammar matchers. People should learn when and how to use those.
- yoz-y 5y agoBut the original SO question does not imply that they want to solve a more complex problem. The SO asker explicitly asked for opinions, so that’s what they got. However, I absolutely think it is the right choice to choose simpler tools to solve simpler problems, as long as you are aware of the implications.
- nixpulvis 5y agoThe regex is surely faster for the specific case. I can't say I've seen an XHTML parser off hand that allows me to stop parsing after just the start tag. Perhaps a lazy parser could start to compete, but I'm just guessing.
- AzzieElbab 5y agoLong regexes are the root of all evil
- crispyambulance 5y agoLong regex's are not the root of all evil but they're certainly a tendril of bad-practice. I would say that the OP's usage, a home-made concoction to find all opening tags in an xhtml doc, crosses the line into bad practice. Get a room and use a parser, people!
- goto11 5y agoSo if using regular expressions is "bad practice", how should one write the tokenizer or lexer stage of a parser?
- crispyambulance 5y agoWho said regular expressions are bad practice? Regexes are great but they can get abused easily. Looking more carefully at the SO question, I am inclined to ask "why?" at least a couple of times because I suspect the answer to the deeper problem the SO OP needs to solve can be worked out with a DOM parser. If not, then definitely a SAX parser could solve that specific problem and it would be more robust than handcrafting a fussy regex.
- goto11 5y agoAs far as I can tell, a SAX parser does not expose the distinction between an opening tag and a self-closing tag. The tag `<foo />` would just emit a startElement event followed by an endElement, exactly the same as `<foo></foo>`. Which means it can't solve the specific problem the OP describes.
- crispyambulance 5y agoI would rather use a robust battle-tested parser and handle that edge case separately than to create regex for all of xhtml, copied off an opinionated blog, as a "hail-mary".
- jancsika 5y agoI love seeing the weirdo CDATA thingy in there! CDATA ftw! E.g., you've got this enormous spec for SVG which includes CSS, but that CSS has syntax inside a style tag which could break XHTML parsers. Amateurs out there are probably thinking, "Well, why not just compromise in the spec and tell implementers to do the same thing that HTML does to parse style tags?" Well, professionals know that cannot work for myriad reasons you can read about if you take out a college loan and remain sedentary for the required duration. The right approach is to throw the CSS stuff inside CDATA tags to tell the parser not to parse it so things don't break. That is the way sensible, educated professionals solve this problem. I'm only kidding! For inline SVGs the HTML5 parser simply says, "Parse this gunk as HTML5, and use sane defaults to interpret the parsed junk in the correct svg namespace so that all the child thingies in that namespace just work." Which it does. Unless you're going to grab the innerHTML of the inline SVG and shove it into a file to be used later as an SVG image. In that case you cross the invisible county line into XHTML territory where the sheriff is waiting to throw you in jail for violating the CDATA rule. In that case the XHTML parser hidden in the guts of the browser doles out the justice of an error in place of your image. Because that is the way sensible, educated professionals solve this problem. :) My holy grail-- how do I use DOM methods to create a CDATA element to shove my style into? If I could know this then I can jump my Dodge Charger back and forth into XHTML without ever getting caught.
- camehere3saydis 5y ago>My holy grail-- how do I use DOM methods to create a CDATA element to shove my style into? If I could know this then I can jump my Dodge Charger back and forth into XHTML without ever getting caught. Does this help? https://developer.mozilla.org/en-US/docs/Web/API/Document/createCDATASection https://developer.mozilla.org/en-US/docs/Web/API/Document/cr...
- jancsika 5y agoAh, thanks! In hindsight I probably could have guessed at "document.create" and then just read the autocomplete suggestions in devTools. :)
- 5y ago
- amelius 5y ago1. determine size of XHTML input 2. build regex that works up to the size determined in step 1. 3. apply regex
- BiteCode_dev 5y agoWeird article that basically says people are wrong then prove they are right.
- inopinatus 5y ago> The question is about finding opening tags in XHTML using a regular expression Bzzzt, wrong, sorry! The question is about finding open tags in the presence of XHTML self-closing tags. That difference alone places these interpretations gulfs apart. But there’s more: it does not specify that the input document is even XHTML, only that XHTML-style self-closing elements may be present. In fact the original question was barely minutes old and tagged merely “regex” when that famous answer was written in 2009; the question was not tagged with “xhtml” until 2012, and not by the original author either. Revealingly, then, if we review the broader context (i.e question history) of the original question author, it’s clear that yes indeed they were trying to fix a malformed document, and in particular to normalise it into XHTML, with focus on fixing up any so-called “dangling tags”. For this task, the suggestion of “use a parser” is indeed sound advice. The real moral here is, don’t be a jerk about the precise semantics of a question, look at what the person needs, and help them ask better questions. Otherwise, you’re just gonna discover that there’s always a bigger jerk, and they’re on Stack Overflow, moderating your stuff.
- swiley 5y agoHere's the original text for reference: Locked. Comments on this question have been disabled, but it is still accepting new answers and other interactions. Learn more. I need to match all of these opening tags: <p> <a href="foo"> But not these: <br /> <hr class="foo" /> I came up with this and wanted to make sure I've got it right. I am only capturing the a-z. <([a-z]+) *[^/]*?> I believe it says: Find a less-than, then Find (and capture) a-z one or more times, then Find zero or more spaces, then Find any character zero or more times, greedy, except /, then Find a greater-than Do I have that right? And more importantly, what do you think? html regex xhtml Share Improve this question edited May 26 '12 at 20:37 community wiki 11 revs, 7 users 58% Jeff
- inopinatus 5y agoNot sure which revision that's pasted from, but it is not the original text, which can be found at https://stackoverflow.com/revisions/1732348/1 https://stackoverflow.com/revisions/1732348/1 and I strongly advise against neglecting the title; doing so is how some folks blunder into misinterpretation, since the wording of that title is telegraphing a quite different underlying need to "please solve this regex puzzle".
- nooyurrsdey 5y agoThis is one of those things that people will debate about endlessly and ultimately it feels so silly. The poster asked how to do it, and this person provided a practical regex to cover most (if not all) cases. Everything else is just pedantic debate.
- h2odragon 5y agoThe pedantic discussions can be fun and educational, but the regex based hack gets the job done in a few minutes, while the pedants are still wrestling with parsers and libraries. ... and then there's the anticipated joy of seeing the pedants' complicated, theoretically correct solution explode because the input wasn't what they assumed, in the first place. The pedants that have that experience either become enlightened, or vociferously strident about the importance of proper, theoretically correct solutions in place of quick hacks. Thus the meme status of the SO answer.
- IshKebab 5y agoThe issue is that sometimes you should use a robust parser and do it properly, and sometimes a hacky regex is fine. But people forget that when arguing about which you should use.
- motoboi 5y agoThe problem is: you cannot parse malformed (real, everyday) html with regexes. But if you need to parse html even malformed) generated by the same template (like a scrapping situation), the whole file becomes regular, which can be parsed by a regular expression. But if you try to parse html in general, too bad because then you’ll need to take html in consideration and will need a recursive descent parser, not a regex. This question popped up so many times in forums in 2000’s that people got mad at that.
- tester756 5y agouhh? just because you can doesn't mean you should just take a look at proposed regex >( > # match all tags in XHTML but capture only opening tags > <!-- .? --> # comment > | <!\[CDATA\[ .? \]\]> # CData section > | <!DOCTYPE ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* > > | <\? .? \?> # xml declaration or processing instruction > | < \w+ ( "" [^""] "" | ' [^']* ' | [^>/'""] )* /> # self-closing tag > | < (?<tag> \w+ ) ( "" [^""]* "" | ' [^']* ' | [^>/'""] )* > # opening tag - captured > | </ \w+ \s* > # end tag > ) it's ugly as hell >Parsing typically uses (at least) two steps: Tokenization which uses regular expressions to splits the input string into a sequence of syntax elements I don't use regex for tokenization, I'm doing something wrong? But overall I think this is important post, even despite I believe that regex is the best example of "good idea, shitty API"
- asddubs 5y agoproblem is that it doesn't reflect how browsers parse things. if you were using this in a security context, e.g., here's an example it won't detect (granted this is not technically valid, but does it matter?): <div "> put arbitrary html here as you please (using single quotes for attributes)<div ">
- SavantIdiot 5y agoOh I've seen this many times in different forms. Especially with regexes. You know what this is a great example of? A case where hacking makes a mess, and thinking before coding solves the problem. The madness comes from using the wrong tool for the problem. Yes, you can hack a regex to parse XHTML this might be "good enough", but it is more robust, cleaner and easier to explain if you use a lexical tokenizer and a grammar. The lure is an illusion that comes from an initial effort assessment. Where the effort to hack a quick-and-dirty regex (call this Ehack) vs a "oh, man, you mean I gotta think about the problem" (call this Ethink) appears as "Equick <<< Ethink." However, it soon evolves to the scenario where "Equick >>> Ethink," driven by the thought process, "I'm almost there, this regex just needs one more tweak." Aka, the gambler's fallacy: it comes into play and the sunk costs are ignored. TL;DR - Use the right tool for the problem, even if it means a slightly larger up-front effort investment.
- Rapzid 5y agoYeah, and you never actually end up solving the problem. You just end up solving every edge case that comes up :D It's the same mentality that can lead to fixing but symptoms without making any real progress on the underlying issues.. Sometimes that's good enough I guess.
- chubot 5y agoThis conversation would be a lot clearer with a distinction between "regexes" and "regular languages". The former is what Perl/Python/etc. have, and the latter is a mathematical concept (and automata-based non-backtracking engines like RE2, re2c, and rust/regex are closer to this set-based definition). https://www.oilshell.org/blog/2020/07/eggex-theory.html https://www.oilshell.org/blog/2020/07/eggex-theory.html With those definitions, this part of the snarky answer is wrong: HTML is not a regular language and hence cannot be parsed by regular expressions That is, regular expressions as found in the wild can parse more than regular languages. (And that does happen to be useful in the HTML case!) This answer is also irrelevant, since the poster is asking for a solution with regexes, NOT regular languages: I think the flaw here is that HTML is a Chomsky Type 2 grammar (context free grammar) and a regular expression is a Chomsky Type 3 grammar (regular grammar). In this post, the example given IS a regex, but it IS NOT a regular language: <!-- .*? --> # comment The nongreedy match of .*? isn't a mathematical construct; it implies a backtracking engine. I gave my analysis here and listed 3 or 4 caveats: https://news.ycombinator.com/item?id=26359556 https://news.ycombinator.com/item?id=26359556 I prefer to use regular languages and an explicit stack, but this is not really what the original question was asking.
- a1369209993 5y ago> This conversation would be a lot clearer with a distinction between "regexes" and "regular languages". Very much so. > In this post, the example given IS a regex, but it IS NOT a regular language: `<!-- .*? --> # comment` The nongreedy match of .*? isn't a mathematical construct; it implies a backtracking engine. Actually, that's [edit: "it IS NOT a regular language"] wrong, at least in principle. If you're limiting it to only the shortest match (which is how HTML (and most other) comments actually work), then (just like `(abc)+` is shorthand for `(abc)(abc)*`) `<!-- .*? -->` is shorthand for (assuming I haven't made a mistake in the for-lack-of-a-better-word-arithmetic): <!-- ([^ ]| [^-]| -[^-]| --[^>])* --> That is, shortest-repetition-only can be implemented in a purely regular system. On the other hand, if you want to allow longer matches when actually needed, then for purely-regular purposes (where it either matches or not) `<!-- .*? -->` is just a wierd way of writing `<!-- .* -->` (which is quite obviously a regular language).
- funyunpowder 5y agobased on the stackoverflow thread, and then the comments here, an interesting research paper topic would be 'why do people get so passionate about regex'
- jll29 5y ago> I think the flaw here is that HTML is a Chomsky Type 2 grammar (context free grammar) and a regular expression is a Chomsky Type 3 grammar (regular grammar). Note that regarding formal language and complexity theory, while it is correct that in general, arbitrary nested structures require a context free grammar (type 2 in the Chomsky hierarchy) and are thus beyond regular (type 3) [1], this statement is NOT true _if_ you limit the nesting depth with a finite constant k. For example, if you agree to an HTML tag maximum nesting depth of, say, 100, then it can be modeled with a regular (type 3) grammar, including correct required matching of opening and closing tags, and hence you can write a regular expression that matches it as well. This debate is well-documented in the theoretical linguistics literature, where some say human languages are not regular because you can always embed yet another additional relative clause in any sentence in principle without adversely affecting grammaticality, whereas others say while you could you won't find natural examples in human-written text documents where extreme nesting depth is actually found. At that point psycholinguists and theoretical linguists usually start a fight about whether memory limits are important or "just performance as opposed to competence". (Goes to show how practical solid theory is.) [1] https://www.sciencedirect.com/science/article/pii/S0019995859903626/pdf?md5=9d466f851651bd592afa5ee561b7a0b0&pid=1-s2.0-S0019995859903626-main.pdf https://www.sciencedirect.com/science/article/pii/S001999585...
- cyberdelica 5y agoIt goes to show, how few people are able to think for themselves. The question originally asked, is "how to match HTML tags". Not how to parse. Not how to scrape. Simply "match". To which I would say, regex is perfectly suited to the task. Furthermore - if one simply needs to scrape content, regex is again, perfectly suited. Scraping, is not parsing - and has no real need for a full blown DOM parsing library. Cargo cult parrots like to say - if the HTML content changes, then one's regex will fail. Well, so will one's DOM parser.