9 ms·
I believe Zalgo has the answer to this, via an equivalent question. https://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-containe
by gamache 7y ago
I believe Zalgo has the answer to this, via an equivalent question. https://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-contained-tags https://stackoverflow.com/questions/1732348/regex-match-open...
- abraCadabstrax 7y agoPerhaps the most epic reply I have ever seen on SO.
- hk__2 7y agoIt’s a famous post on StackOverflow, but I don’t find it particularly helpful.
- joppy 7y agoI don't think that answer was written with the intent of being particularly helpful, I think it was written with a different goal in mind.
- jacobush 7y agoIt is helpful in that almost koan way though.
- swasheck 7y agoIt is helpful. It may not provide the answer the OP wanted, but it provides the answer they needed.
- wool_gather 7y agoThe author said at one point that it was written in a bit of a huff of frustration, and if I remember correctly, also after a pint or two. It's not the best answer answer by any means, but it is one of those small gems of the web, I think.
- snazz 7y agoHad it not gotten extremely lucky, it would have a very negative score. Anything intended to be funny on StackOverflow doesn’t go over well very often.
- ChrisSD 7y agoThe answer was made in '09. StackOverflow was more accepting of some humour at the time. And it is making a valid point: regex isn't the right tool for this job.
- unlinkr 7y agoThe question is about how to identify start and end-tags in XHTML. What would be an appropriate tool for that job?
- shakna 7y agoA parser. Specifically, an XHTML parser.
- unlinkr 7y agoHow do you think an XHTML parser is written? In particular, how does an XHTML parser identify tokens like start and end tags?
- saagarjha 7y agoIt keeps track of state that a regular expression cannot?
- unlinkr 7y agoYou don't need to keep track of state to match tokens like XHTML start or end tags.
- journalctl 7y ago
- stfwn 7y agoIf you were looking for the reason why a regex cannot parse HTML, it is because HTML has matching nested tags and regex parsers are finite state machines (FSM). What this means is that a regex parser is like a goldfish. It only knows about the state it is currently in (what it just read) and which possible states it may transition to (what is legally allowed to come next). The fish never remembers where it was before; there is no option to have the legality of a transition depend on what it read _before_ the current state. But this memory function is a requirement to recursively match opening tags to their closing tags - you need to keep a stack of opening stacks somewhere in order to then cross off their closing tags in reverse order. So regex cannot parse HTML.
- unlinkr 7y agoThe question is about identifying end-tags in XHTML. This is indeed possible with a regex.
- bryanrasmussen 7y agotheoretically I believe an end tag really requires a valid start tag. anyway you can probably answer any number of simple questions about a bit of HTML using regex but as code wants to grow to handle more use cases there will come a time when the solution will break down and the code that wrote to handle all the previous uses will need to be rewritten using something other than regex.
- unlinkr 7y agoAn element requires a start and end tag, or a self-closing start tag.
- unlinkr 7y agoYou have to distinguish between the different levels of parsing. Regexes are appropriate for tokenization, which is the task of recognizing lexical units like start tags, end tags, comments and so on. The SO question is about selecting such tokens, so this can be solved with a regex. If you have more complex use cases like matching start tags to end tags, you might need a proper parser on top. But you still need tokenization as a stage in that parser! I don't see what you would gain by using something other than regexes for tokenization? I guess in some extreme cases a hand written lexer could be more performant, but in the typical case a regex engine would probably be a lot faster than the alternatives and certainly more maintainable. I know it is possible to write a parser without a clear tokenization/parsing separation - but it is not clear to me this would be beneficial in any way.
- kortex 7y agoI found it extremely helpful. The sheer emphasis of the reply made me very curious why the idea of using regex on xml is so bonkers. - I will never forget that regex can't parse XHTML - the reason being, regex is insufficiently powerful - when I first saw this post, I knew little about regex under the hood, this sent me down a wiki hole of FSMs, pushdown automata and turing machines - this misconception is apparently common enough to be madness-inducing to those that know better - use a hecking xml parser instead It almost reminds me of a Bill Nye sketch. Teaching through a bit of non-sequitur and absurdism.
- unlinkr 7y agoThe problem with the answer is it is wrong. The question is about identifying start-tags in XHTML. This is a question of tokenization and can be solved with a regular expression. Indeed, most parsers use regular expressions for the tokenization stage. It is exactly the right tool for the job! Furthermore, the asker specifically needs to distinguish between start tags and self-closing start tags. This is a token-level difference which is typically not exposed by XHTML parsers. So saying "use a parser" is less than helpful. I have elaborated a bit in blog post: https://www.cargocultcode.com/solving-the-zalgo-regex/ https://www.cargocultcode.com/solving-the-zalgo-regex/
- LoSboccacc 7y agoit's sad that today the stack overflow majority shares your mindset and an answer like that would drown in downvotes or killed by moderation. some people like to act super serious all the time like they're playing a sitcom version of what they think adulthood is in a quest to be the most boring person on earth like if that's the goal of human interaction
- hk__2 7y agoYou can be funny and helpful at the same time.
- LoSboccacc 7y agowhich is exactly what that so answer was
- unlinkr 7y agoThe problem is it is funny and wrong. Apparently it have given a lot of people really confused ideas about what is possible and what is not possible with regular expressions. If it had been funny and right I would not have a problem with it.
- LoSboccacc 7y agoStack overflow is not a homework help group, the "Regex are the wrong tool for the job" answer is more correct than the answer with the working Regex.
- unlinkr 7y agoThe job (the question asked) is about recognizing start and end tags in XHTML. These are lexical tokens and therefore regular expressions are a perfectly fine tool for this. Indeed many parsers use regular expressions to perform the lexical analysis (https://en.wikipedia.org/wiki/Lexical_analysis https://en.wikipedia.org/wiki/Lexical_analysis) aka tokeniziation. Quoting from wikipedia: The specification of a programming language often includes a set of rules, the lexical grammar, which defines the lexical syntax. The lexical syntax is usually a regular language, with the grammar rules consisting of regular expressions; If you disagree that regular expressions are an appropriate tool for lexical analysis, can you articulate why? And what technique do you propose instead?
- hammerbrostime 7y agoWow it looks like It was designed by David Carson.
- nolok 7y agoIn SO's hall of fame of regular members being feed up of a common yet stupid question always coming back up it's up there alongside the thread about making addition with jQuery
- lolinder 7y agoMy favorite part of it is the moderators' note at the end: > Moderator's Note > This post is locked to prevent inappropriate edits to its content. The post looks exactly as it is supposed to look - there are no problems with its content. Please do not flag it for our attention.
- gnud 7y agoThe big problem with the content is suggesting to use an XML parser to parse HTML. That might often work, and several XML parsers have some sort of HTML "mode", but in general, no, not really. HTML documents are frequently invalid XML. HTML documents lack an XML declaration, they don't close all tags, etc. They might even lack a root element. Someone mentioned this in a comment on SO, but it was (as far as I could see) not really addressed or answered. HTML 5 has it's own parsing rules, specified at [1]. However, I don't know of any implementation of these outside of browsers. I normally use Beatuiful Soup [2] (for Python) or HTML Agility Pack [3] (for .NET), and while I don't think these implement the exact standard, they're easier than struggling with a XML parser for HTML "in the wild". 1: https://html.spec.whatwg.org/multipage/parsing.html https://html.spec.whatwg.org/multipage/parsing.html 2: https://www.crummy.com/software/BeautifulSoup/ https://www.crummy.com/software/BeautifulSoup/ 3: https://html-agility-pack.net/ https://html-agility-pack.net/
- kortex 7y agoThe question title implies he's parsing XHTML text. > XHTML is an XML-based HTML. It serves the same function as HTML, but with the same rules as XML documents. https://stackoverflow.com/q/1429065 https://stackoverflow.com/q/1429065
- ChrisSD 7y agoThe question title is ambiguous but the question tags include both HTML and XHTML. Self-closing `<element />` is an XML-ism which in HTML is treated the same as `<element>` but it was still fashionable to use them in HTML when XHTML was at its height.
- unlinkr 7y agoHere is the answer to the question: https://www.cargocultcode.com/solving-the-zalgo-regex/ https://www.cargocultcode.com/solving-the-zalgo-regex/ tl;dr: It can indeed be solved relatively easily with a regex.
- kortex 7y agoThis is a bit out of my wheelhouse, but this feels wrong, or at least naively capable. Like it feels like this sort of reasoning leads to the kind of bugs (depending on what you use the result of the rexex for) that allow for code injection, a la the Equifax hack. Maybe another HN poster can back me up, or explain why in fact Zalgo is mistaken and CargoCode is correct. Either way, this sort of complexity is one reason I avoid XML like the plague and keep HTML at arm's length.
- lonelappde 7y agoCargoCode is correct. Zalgo simply misread the question because he was so sick of similar subtly different questions.
- unlinkr 7y agoThis is what I really hate about the Zalgo answer. It is instilling people some vague sense that regular expressions are somehow bad, wrong and dangerous. But without any real arguments or contexts which would allow you to evaluate if the feeling is justified.
- lol768 7y agoIt doesn't work for me with regex101. "The preceding token is not quantifiable" on this part: | < (? \w+ )
- kortex 7y agoSee, this is kinda what I mean. Maybe you can detect tags with regex, but maybe you shouldn't, given the widespread but subtle differences in regex engines. Perhaps the entire approach of "why are you trying to parse X?" Needs to be traced and re-evaluated.
- tzs 7y agoFrom that answer: "You can't parse [X]HTML with regex". I've always wondered how that came about. How come very early in the development of the web no one important enough for people to pay attention to them said, "Hey...wait a second. If this thing becomes popular, people are going to really want to processes web documents with their usual text file processing tools and techniques. We really outta make this thing reasonably easy to process with grep and sed and awk and Perl and such"?
- tom_mellior 7y agoThe problem isn't with the surface syntax of HTML. Anything else, like S-expressions or JSON or whatever would have the same problem. The problem is arbitrary nesting of elements. No language with nesting can be parsed with regex (modulo various copouts mentioned elsethread). But such nesting is useful and necessary for a document language like HTML. This goal is important; "we must be able to use awk" isn't.
- deleted 7y ago[deleted]