4 ms·
Why you should not parse (X)HTML with a Regexp
- iwwr 16y agoEven Jon Skeet cannot parse HTML using regular expressions.
- telemachos 16y agoBut according to a comment on the post, "Chuck Norris can parse HTML with regex" (emphasis in original). More seriously, when SO was newer, I remember the feeling that these things came in waves. For a few months, there was someone who responded to nearly every offending Ruby question (or answer) by pointing out that exceptions shouldn't be used for flow control. See these two by another poster for more on HTML and regexes[1][2]. [1] http://stackoverflow.com/questions/701166/can-you-provide-some-examples-of-why-it-is-hard-to-parse-xml-and-html-with-a-rege http://stackoverflow.com/questions/701166/can-you-provide-so... [2] http://stackoverflow.com/questions/773340/can-you-provide-an-example-of-parsing-html-with-your-favorite-parser http://stackoverflow.com/questions/773340/can-you-provide-an...
- obtino 16y agoThe technical explanation for this is given in comment 3 of the page and sums it up perfectly: "I think the flaw here is that HTML is a Chomsky Type 2 grammar (context free grammar) and RegEx is a Chomsky Type 3 grammar (regular expression). Since a Type 2 grammar is fundamentally more complex than a Type 3 grammar - you can't possibly hope to make this work. But many will try, some will claim success and others will find the fault and totally mess you up." More info: http://en.wikipedia.org/wiki/Chomsky_hierarchy http://en.wikipedia.org/wiki/Chomsky_hierarchy
- wvl 16y agoAnd the previous discussion: http://news.ycombinator.com/item?id=1487695 http://news.ycombinator.com/item?id=1487695
- d_r 16y agoFortunately, BeautifulSoup saves the day for HTML parsing tasks. (http://www.crummy.com/software/BeautifulSoup/ http://www.crummy.com/software/BeautifulSoup/)