4 ms·
That is not XHTML.
by unlinkr 7y ago
That is not XHTML.
- nurettin 7y ago<input value="how about this? />"/>
- unlinkr 7y agoThat is a valid XHTML tag (if I remember correctly) and can be matched perfectly fine by a regex.
- nurettin 7y agoPerhaps something like "([^"]*)" could skip what is inside the string literal. Unless there is "<input" in the string literal, then where you start parsing becomes very important.
- unlinkr 7y agoThat pattern would indeed match a quoted string. I don't see how it would matter if the quoted string contains something like "<input". It can contain anything except a quote character.
- hk__2 7y agoWhat about <a>this <!-- </a> --> </a> <!-- </a> -->?
- unlinkr 7y agoYes you can tokenize this with a regular expression and extract the valid start and end tags. If comments in XHTML could nest you would have a problem. But this is not the case.
- hk__2 7y ago> Yes you can tokenize this with a regular expression and extract the valid start and end tags. So you need more than a regular expression, hence your premise is incorrect.
- unlinkr 7y agoNo, you don't need more than a regular expression. If you want to extract elements, i.e. match start tags to the corresponding end tags, then you need a stack-based parser. But just to extract the start tags (which is the question) a regular expression is sufficient. The original question is a question about tokenization, not parsing, which is why a regular expression is sufficient.