4 ms·
Step 1) Install pyre2 2) import re2 as re 3) 10x faster regex
- sandGorgon 16y agoCould this be uploaded to pypi, so that it is available via pip ? pip install pyre2 -E /tmp/testing
- maxiak 16y agoIf you read the readme under installation, you'll find that it is. pip install re2 -E /tmp/testing
- natmaster 16y agoWhat kind of licensing does this have? It would be interesting to see if this could make it in Python 3.2 (It's supposed to include the Unladen Swallow stuff, so might as well take more awesomeness from Google.)
- mcav 16y agoThe 3-clause BSD license... http://github.com/axiak/pyre2/blob/master/LICENSE http://github.com/axiak/pyre2/blob/master/LICENSE
- axiak 16y agoBSD (as does the google RE2 C++ module)
- dododo 16y agonote it depends upon cython: i suppose it's more likely it will make it into cython rather than python 3.2.
- axiak 16y agoLike the libxml module which uses Cython, the compiled CPP module is distributed along with it, so you don't actually need Cython to install it or compile it. But you would need Cython to do any development on it.
- kingkilr 16y agoIt won't. It doesn't implement the entire regular expression language Python supports.
- kvs 16y agoPlease note that RE2 doesn't support backtracking.
- aaronblohowiak 16y agoif you're using backtracking, maybe "regular expressions" aren't the right tool.
- dkarl 16y agoCould a regex library parse a regular expression and decide reasonably quickly which engine to use?
- kingkilr 16y agoProbably, but python-core has no interest in maintaining multiple regular expression implementations: http://mail.python.org/pipermail/python-dev/2010-July/101606.html http://mail.python.org/pipermail/python-dev/2010-July/101606... (this thread starts talking about regex, which is a backwards compatible enhancement to re, but also covers re2)
- logic 16y agoIf you're not familiar with re2, you may be familiar with a couple of other little projects that the author, Russ Cox, is involved with: Go (http://golang.org/ http://golang.org/) and Plan 9 (http://swtch.com/plan9port/ http://swtch.com/plan9port/). Also, here's a great bit of technical history behind Russ' re2 implementation, and why pyre2 will never be completely compatible with Python's re (without fallback): http://swtch.com/~rsc/regexp/regexp1.html http://swtch.com/~rsc/regexp/regexp1.html http://swtch.com/~rsc/regexp/regexp2.html http://swtch.com/~rsc/regexp/regexp2.html http://swtch.com/~rsc/regexp/regexp3.html http://swtch.com/~rsc/regexp/regexp3.html
- pfedor 16y agoGoogle Code Search was his intern project: http://googlecode.blogspot.com/2006/10/search-worlds-public-source-code.html http://googlecode.blogspot.com/2006/10/search-worlds-public-...
- ww520 16y agoI'm curious. How much of a speedup in regex when using DFA instead of NFA? I believe most regex implementation use NFA but there are well known algorithms to convert NFA to DFA. It should be worth the try.
- ori_b 16y agoThe problem is that in the worst case, converting an NFA to a DFA gives an overhead of 2^n, while keeping it as an NFA only incurs an overhead of m (where m is the regex length). So, in the worst case (ie, under algorithmic complexity attacks), the NFA implementation is O(m * n), while the complexity of the DFA implementation is - I believe - O(2^m * n)
- koenigdavidmj 16y agoAnd fortunately, if re2 can not handle one of the regexen that you throw at it, then it will pass it down to Python's original re module (NFA_based). Perfectly seamless.
- ori_b 16y agoBoth re2 and Python are NFA based. The difference, I believe, is that re2 doesn't support backtracking, which changes the worst-case time from linear growth to exponential growth on the input length. (For the snobs, that means that re2 regexes are regular expressions. Python's aren't.)
- deleted 16y ago[deleted]
- democritus 16y agoI tried using pyre2 and experienced runaway memory usage. I was using re.split. The behaviour went away when I switched to the standard re module. (without changing my code.) Did anyone else experience this? The code in question was: (a, b) = re.split("\n\s*\n", text, maxsplit=1) Thank you for any insights you may have regarding this.