5 ms·
You might realise that parsing HTML or C with regular expressions is a bad idea.
by lindig 13y ago
You might realise that parsing HTML or C with regular expressions is a bad idea.
- laureny 13y agoReally? http://stackoverflow.com/a/1732454/162410 http://stackoverflow.com/a/1732454/162410
- UNIXgod 13y agoIt's tricky with HTML but can be done. The C language was one of it's initial use cases.
- theseoafs 13y agoNo, it can't. As in, you cannot parse HTML or C with regular expressions, and you can prove that that's the case. Regexes simply aren't powerful enough. This is one of the advantages of studying formal languages -- you get to learn about really useful models of computation (here, DFA's), but you also get to learn about their limits.
- dllthomas 13y agoScanning the C language is done using regular expressions, typically. Parsing the resulting token stream cannot be done with only regular expressions. At least not with the theoretical construct. Some "regular expressions" libraries, notably PCRE, let you do plenty that goes beyond the theory at the expense of some efficiency. Parsing some context-free grammars with these libraries is certainly possible, but you're not really restricting yourself to "just" regular expressions at that point. Regular expressions plus interwoven use of a stack is a pushdown-automata, which can totally parse C and HTML.
- dllthomas 13y agoParsing C or HTML with regular expressions is a great idea. Parsing C or HTML with only regular expressions isn't.