4 ms·
Yes, this ^ Because, no, you can't parse XHTML with regex. As easily shown by the pumping lemma and all that jazz. But, there's no freaking reason why you can
by pidge 14y ago
Yes, this ^
Because, no, you can't parse XHTML with regex. As easily shown by the pumping lemma and all that jazz.
But, there's no freaking reason why you can't tokenize an XML start tag with a regex! In fact, you'll probably find that most uses of parsers in real life have regexes to tokenize down at the level that they can handle, before using a parser on the resulting tokens for the part that actually needs to be a CFG (among other reasons, because a compiled FSM is a lot faster than even a limited LALR parser).
Looking at this specific example, we can refer to the definitions for start tags [1] and empty element tags [2] in XML, and see that all their constituent rules form a regular language (if you don't believe me, it's not too hard to go check for yourself). So, especially since the orignal question doesn't even mention 'parsing', can we all please just shut up? (unless you actually want to figure out the horrible mess necessary to define a regex from the spec :P )
1. http://www.w3.org/TR/xml11/#sec-starttags http://www.w3.org/TR/xml11/#sec-starttags
2. http://www.w3.org/TR/xml11/#dt-eetag http://www.w3.org/TR/xml11/#dt-eetag