9 ms·
Why Using .* in Regular Expressions Is Almost Never What You Actually Want
- blt 12y agoVery interesting. I've had the greedy .* "overmatch", like probably almost everyone who's used regular expressions. Had no idea they are a performance drain even when giving the right answer though. I like posts about details of software craftsmanship like this.
- prohor 12y agoA bit of a problem with lazy quantifiers is that they are not so widely supported out of the perl world. Therefore I often need to find some tricks to get similar behavior (eg. "[^,]*," - if coma is separator)
- BugBrother 12y agoHuh, PCRE is universal (almost, not in e.g. Emacs :-( ) and supports (almost) everything? Edit: Cough, after rechecking... PCRE is not as universal as I thought. It seems I've been lucky. :-) http://en.wikipedia.org/wiki/Comparison_of_regular_expression_engines#Part_2 http://en.wikipedia.org/wiki/Comparison_of_regular_expressio... Edit 2: "Atomic groups" on that wikipedia link is when you can write a full grammar in a large regexp, right? Answer myself: No, it is the name for stopping backtracking. I've seen it as named "possessive" (perldoc perlre).
- collyw 12y agoMySQL doesn't have them, it has some POSIX version, which is way less intuitive. Plenty of people use MySQL.
- BugBrother 12y agoIsn't that part of an SQL standard (1999?).
- SiVal 12y agoI thought it was becoming more universal, but now I'm not so sure. grep on the Mac used to use PCRE regexes if you used the -P option (`grep -P ....`), but beginning with OS X 10.8 the -P option was removed, so an important place that used to offer PCRE (default grep on a default Mac) actually removed support for it. They didn't replace it with something better; they just took it away. Maybe a Unicode issue? Not only is PCRE not universal, overall support might even be waning.
- hnriot 12y agoExcept they are! Java, python, JavaScript etc all support lazy quantifies. In fact I can't think of a single language that doesn't.
- perlgeek 12y agoBasically the same advice, from the year 2000: http://www.perlmonks.org/?node_id=24640 http://www.perlmonks.org/?node_id=24640 There is one legitimate use case of .* though: advancing to the last match of something. If you want to find the last digit in a string, /.*(\d)/ will readily find it for you.
- Mithaldu 12y agoBe aware though that \d will also find arabic or roman numerals. :)
- mhaymo 12y agoRoman numerals? I have never encountered a regex implementation where '\d' matches 'X'. It's certainly not the case for Javascript: http://regexpal.com/?flags=g®ex=\d%2B&input=10%0Ax%0AX%0AVII%0AMMXIV http://regexpal.com/?flags=g®ex=\d%2B&input=10%0Ax%0AX%0A...
- mmastrac 12y agoThere's a whole lot of these numeral characters in unicode, for example: http://www.charbase.com/2169-unicode-roman-numeral-ten http://www.charbase.com/2169-unicode-roman-numeral-ten (even more in http://www.charbase.com/block/number-forms http://www.charbase.com/block/number-forms) I'm not sure if Javascript matches these in its \d pattern, however, but I think that most regexp engines default to the ascii [0-9] unless you are using \p{Number}.
- fennecfoxen 12y ago> /\d/.test("\u2169") false
- pyre 12y agoThere's also the character classes in Unicode: http://www.fileformat.info/info/unicode/category/index.htm http://www.fileformat.info/info/unicode/category/index.htm
- gesman 12y ago.*? Is the solution!
- deleted 12y ago[deleted]
- tjgq 12y agoAlso read Russ Cox's writeup on implementing regular expressions [0]. Backtracking can be done efficiently; it's just that most regular expression engines have suboptimal implementations for it. [0] http://swtch.com/~rsc/regexp/regexp1.html http://swtch.com/~rsc/regexp/regexp1.html
- BugBrother 12y agoI've cursed over Python's backtracking, at least a few years back. (Why can't they just use PCRE? :-( Any advantage at all?)
- ori_b 12y agoPCRE has the same problems with backtracking.
- BugBrother 12y agoAs bad? I might have had smarter coworkers at different times... :-)
- maxerickson 12y agoPython beat PCRE to Unicode support by several years. (So there at least was an advantage)
- BugBrother 12y agoI thought the Unicode support was still spotty (< 3.X)? Or you mean the support is better than in PCRE?
- maxerickson 12y agoIn 2000, PCRE simply didn't support Unicode. Python 1.6 and 2.0 did (at least, based on some quick searching PCRE added support for Unicode in 2004). "spotty" probably isn't the right word either, the change in 3.0 was to default to treating text as always being Unicode, the 'unicode' type in 2.x is reasonably complete (as these things go), just not the default treatment for text.
- guynamedloren 12y agoI've run into this problem so many times. Everywhere I think I want .* , I actually want .*? (non-greedy matching). Make a mental note of this. It'll save you lots of headaches.
- moron4hire 12y agoI think the advent of automatic regex match highlighting in text editors is changing the regex use-case for a lot of people. It certainly did for me. I no longer see regexs as just "something you use in code to test input". I now use them as general purpose text editing tools. In a way, it's like templated text output, with input specified in the same buffer. I know this has been done forever, but usually only by extreme greybeards in Vi or Emacs world. The auto-highlighting now makes it possible for everyone to do it. So that said, with the ability to restrict regexes to just a selection of text, it's more about regex golf--the fewest characters, the most productive--than it is about semantic correctness. If it works for my input, that's all that matters, because the regex is getting discarded thereafter.
- Pxtl 12y agoYeah, I do all my data-imports from flat files using regex - easy to export from spreadsheet programs as flat files, then regex them into a bunch of insert/update statements.
- collyw 12y agoYes, I do plenty of similar things. Format a load of data using regexes first, then use it hard coded as string to do quick one off script to update the database. It beats trying to parse Excel directly, as you never know what data type a cell will return.
- kstenerud 12y agoActually, it only needs to be \[([^,]+),([^\]]+)\] because you're only going up to a comma in the first capture group and a square bracket in the second.
- Pxtl 12y agoI've gotten into the habit of using the "not" operation instead of .* a lot. If I'm looking for bracketed text, I use not-bracket to match the contents. I tend to avoid the non-greedy operator just because it often fails in terrible half-assed regex implementations (eg. visual studio 2010)
- bane 12y agoI wish the not operator allowed for sub-expressions instead of just character classes. It'll probably make it slower, but it would remove lots of unreadable convolutions people have to go through.
- ori_b 12y agoThere are some regex implementations that allow it, but it's a very confusing feature. Remember that '' is not 'a'. Arbitrary expressions can have arbitrary length, so excluding an expression simply will match it, fail the match, and backtrack to the next option.
- blueblob 12y agoMe too. I find that if you work in a bunch of different languages it seems more portable (and one less thing to remember). It also seems easier to debug.
- mschuster91 12y agoFor the "greedy" behaviour, PHP has the "U" flag... dunno about other implementations though.
- chernevik 12y agoTo be picky, it's always what I want, but with a lot of other stuff I don't.
- larubbio 12y agoIs the early example in the document correct? Using an input string of abc123 he claims [a-z]+\d+ will match the entire string (which I agree with). He then says that [a-z]+?\d+? will only match abc1. Wouldn't it fail since the non-greedy match on [a-z] would just match 'a' causing the non-greedy match on \d to fail trying to match 'b'?
- simcop2387 12y agono it'll still match but because both are non-greedy it could match on just c1 instead of abc123.
- prawks 12y agoIt could match on c1, but I believe since most (all?) regex parsers parse left-to-right, it will match the a, look for another a-z character or a digit, find b, repeat, find c, then find 1 which completes the pattern.
- gatehouse 12y agohttp://regex101.com/r/aR5xM2 http://regex101.com/r/aR5xM2 I used this tester posted elsewhere in the thread, it seems like since the lazy components expand "as needed" to achieve a match, it will succeed on "abc1". EDIT: I wrapped it in a group for clarity.
- NanoWar 12y agoThe regex fiddle is really useful: http://regex101.com/r/qQ2dE4 http://regex101.com/r/qQ2dE4
- icambron 12y agoAbout eight years ago I finally read Friedl's Mastering Regular Expressions [1]. I know, right? A 500-page book about regular expressions, a tool I already knew (or thought I did). But it's actually a great book-- easy to read and full of genuinely good information on the how and why of regex, and it totally changed my understanding of them. If absolutely anything in this article surprised you, I highly recommend you read the book. [1] http://regex.info/book.html http://regex.info/book.html
- bthornbury 12y agoI once used .* in a crawler. Came back the next day to find much rogue html amongst the content of my site. I find something like [^
- bthornbury 12y agoI once used .* in a crawler. Came back the next day to find much rogue html amongst the content of my site. I find something like [^{{delimiting character}}]* to be better
- spb 12y agoThis is why I like Lua's '-' for non-greedy matching in its pattern facilities.