6 ms·
I do not know if it's correlation or causation, but almost all code using regular expressions I encounter is horribly broken. 1. Regular expressions are often
by banthar 11y ago
I do not know if it's correlation or causation, but almost all code using regular expressions I encounter is horribly broken.
1. Regular expressions are often used instead of existing parsers. XML, CSV, file paths, URIs etc. all already have fast, well tested and correct parsers.
2. Often the thing worked on should not be a string in the first place. For example comma separated string is used instead of a list and regular expression emulate list operations.
3. Some things could be processed much easier with different tool - for example recursive descent parser. Yet the developer still tries to parse arithmetic expressions with a hammer.
4. They are often developed via trial and error. They are either first thing which worked for a simple case or are 10 line monsters riddled with exceptions from exceptions.
5. They are often part of hacks and workarounds. For example User-Agent is matched to work around bugs in browsers.
3. They give very limited feedback to user. There is either a match or no match. There is no way to tell what and where is broken.
There are valid use cases for regular expressions. It's even possible to write correct code with them. It's just a rare sight.
- ivanhoe 11y agoMany of these points are correct, but a little comment on #1: Often you don't care about the whole structure, you just need some small piece of data from a middle of a huge document. One common example is a spider that collects a price of some product from a 200+KB webpage. You just need those few digits and don't care about the head or title or the structure of dom or anything else. In such cases (and it's very common task for people working on data extractions) no complex parsers can ever compete with the regexp in terms of speed and memory footprint. And if you need to parse a few millions of products that performance gain is a huge deal. So don't underestimate the power of regexp when properly used...
- banthar 11y agoThat is fine for throwaway scripts. But such "perl duct tape" is not 100% accurate and will break for no reason. There is no place for such solutions in reliable and maintainable software.
- ivanhoe 11y agoWhy would it "break for no reason"?! For all I know regExps matching one small piece of the page if far less prone to breaking than parser that has to analyze the whole page. Designer changes one <div> or id/class somewhere in the top of the DOM tree and you can't reach the node that you are looking for anymore. Same goes for regExp of course, but it's looking at a smaller portion of the html, so it's less likely to be affected by small changes in some unrelated part of the page. And any major redesign will break any dedicated scraper, no matter which parser it uses...
- banthar 11y agoLets try an example. Extract first link address from https://news.ycombinator.com/ https://news.ycombinator.com/. As DOM query: document.getElementsByClassName("title")[0].parentElement.getElementsByTagName("a")[1].href This will break: * When title element no longer has "title" class. * When title is no longer a sibling of link. * When link is no longer 2nd link of its parent. As regular expression: document.documentElement.innerHTML.match('td class="title">.*a href="([^"]*)"')[1] This will break: * On any white space change. * On any new attributes on td or a. * When ' is used instead of " * When href includes escaped " * In most cases when DOM query will break. Many of those can happen without any server-side changes. It will sometimes works sometimes won't - making it hard to test. There are cases when regular expression will break less often than DOM but DOM is easier to reason about, more predictable and has less corner cases.
- kwhitefoot 11y agoAnd the xml paresr will fail if the xml is not well formed while the regex will just keep sailing along. I had an example of this with BlogPoster.py which uses python xmlrpc. There are Wordpress hosts which return invalid xml and this causes an exception, I reimplemented what I needed with Bash and cURL using regexes and it works fine.
- banthar 11y agoI'm not sure if you mean that pro or against regular expressions. That is a good example for #5. Regular expressions are great to quickly patch together something that kinda works, but: 1. Those hosts are still broken. The next person will have to jump the same hoops to support them. 2. You parser is very permissive. It will encourage people to create even more broken implementations. 3. The specification of this protocol is now worthless. There is no way to safely add new functionality. Any new element or attribute can break those regexes. Everyone has to take every implementation into account. 4. You are probably missing some corner cases like CDATA elements or quoted characters. From your point of view, it probably makes sense to support even broken sites. But, you are helping to create next HTML - where every implementation works differently and you have to test everything on every browser.