10 ms·
Python's Hidden Regular Expression Gems
- ioquatix 11y agoThere is nothing unique about Python's implementation of regular expressions. Ruby has an equally (if not more) powerful `StringScanner`. This is a nice article but it could have done without the "it's one of the best of all dynamic languages I would argue" tone.
- deleted 11y ago[deleted]
- lazyjones 11y agoPersonality cult? Perl programmers can only shake their heads at the post, since its regexp implementation is much more powerful and documented properly.
- ioquatix 11y agoI was going to mention Perl, but I decided it goes without saying :)
- simgidacav 11y agoModest and good perl programmer spotted.
- sdoering 11y ago@lazyjones reminded me of this: https://xkcd.com/1090/ https://xkcd.com/1090/
- the_mitsuhiko 11y ago> Ruby has an equally (if not more) powerful `StringScanner` The string scanner is just a step by step matching of individual expressions. Because they are not folded together the "skip non matching" part has been done in Ruby which is precisely what the Python scanner avoids.
- willvarfar 11y agoAnother very-cool undocumented feature on another regex engine is re2's Set. It compiles a collection of regex to a single regex, and so allows you to very efficiently match a string against an array of patterns.
- andreasvc 11y agoWhich re2? I maintain a fork of re2 but it's not in there [1]. If you mention re2 the main cool feature about it is that it is efficient, matching in linear time using DFA. Unfortunately unicode strings need to be encoded to utf8 but if you can design your application to work with utf8 bytestrings you can avoid that cost. [1] http://github.com/andreasvc/pyre2 http://github.com/andreasvc/pyre2
- willvarfar 11y agohttps://github.com/google/re2/blob/master/re2/set.h https://github.com/google/re2/blob/master/re2/set.h <-- c++ API https://github.com/google/re2/blob/master/re2/prog.h#L339 https://github.com/google/re2/blob/master/re2/prog.h#L339 <-- c
- andreasvc 11y agoNeat. I should consider wrapping that.
- fanf2 11y agoAnother package that does this is re2c [1] which is used by spamassassin to compile its rulesets to scan messages faster. [1] http://re2c.org http://re2c.org
- fndrplayer13 11y agoReally interesting read. To be honest, I got a little lost around the Scanner implementation portion. Guess its time to take that example and play around with it myself. My one suggestion would be to maybe walk through an example of how the Scanner would work to demonstrate your point. Thanks for the great post.
- the_mitsuhiko 11y agoI made a larger class in a github repo and an example of what you can do with it here: https://github.com/mitsuhiko/python-regex-scanner/blob/master/examples/wiki.py https://github.com/mitsuhiko/python-regex-scanner/blob/maste...
- fulafel 11y agoI did a double take upon seeing "The regex module in Python is really old by now" and had to make sure it's talking about the current re module! re was introduced alongside the older regex module around Python 1.5, the latter was finally removed in Python 2.5.
- andreasvc 11y agoAdditionally there's a newer module also named 'regex': https://pypi.python.org/pypi/regex https://pypi.python.org/pypi/regex
- rspeer 11y agoAnd this newer 'regex' is actually really good at tricky cases such as matching word boundaries. (An apostrophe or a non-ASCII character is not necessarily a word boundary!)
- andreasvc 11y ago> An apostrophe or a non-ASCII character is not necessarily a word boundary! I don't see how a regular expression library could help with that (other than proper Unicode support), because word boundaries are a language-specific, linguistic problem; i.e., you will need to supply a list of possible contractions anyway. Tokenization of natural language text may appear like a straightforward and solved problem, but there are actually lots of messy details to get right.
- mkesper 11y agoLately there had been a comparison of Javascript, Perl and Python regex machinery posted here iirc (can't find right now) and the external regex module was found to be much better regarding unicode support.
- berntb 11y ago(How could non-3 Python have good Unicode support? :-) ) The default Python regexp gets into a tailspin on less well formulated regexps, which did work well with both PCRE and Perl 5. (I wrote an application (specialized query language) a few years ago where programmers entered regexps. Sometimes the execution just hanged. This was 2.6 and early 2.7.) I haven't tried the external regexp module, but hope it is better.
- lugus35 11y agohttp://doc.perl6.org/language/regexes http://doc.perl6.org/language/regexes Read. Recite. Review.
- andreasvc 11y agoI wonder what the reason is to include code in a release without documenting it. Maybe this article can form the basis for finally documenting this feature? There's also the reverse with Python: useful code in the documentation not included in the standard library.
- HerpDerpLerp 11y agoMaybe this sort of thing will be tackled by the new stackoverflow documentation thing. This link may do nothing if you are not part of the beta! http://docs-beta.stackexchange.com/documentation http://docs-beta.stackexchange.com/documentation
- creshal 11y ago> I wonder what the reason is to include code in a release without documenting it. Presumably it's considered an implementation detail and the author didn't realize it was useful for anything else.
- dalke 11y agoThe author knows that it's useful. See http://effbot.org/zone/xml-scanner.htm http://effbot.org/zone/xml-scanner.htm where the author uses the scanner: > The 2.0 engine provides another (undocumented) feature that can be used to optimize this even further. The scanner method is used to create a scanner object and attach it to a string. See http://bugs.python.org/issue5337 http://bugs.python.org/issue5337 where the Python developers wondered if it should be documented. There's a 10 year old notice at https://mail.python.org/pipermail/patches/2006-November/021090.html https://mail.python.org/pipermail/patches/2006-November/0210... saying that the code would crash if the scanner was called from multiple threads. There's no doubt much more information about this - the above was all I cared to find out.
- notzorbo2 11y agoThis happens pretty often in the Python world. There's a bit of an unwritten rule to leave implementation details public that would be private in other languages. Many libraries simply don't bother with prefixing privates with '_' and just leave things undocumented that you probably shouldn't touch/use. One notable example is importing libraries in your code automatically exposes them to the caller. $ cat lib.py import re def somefunc(): pass $ python >>> import lib >>> lib.re <module 're' from '/usr/lib/python2.7/re.pyc'> Many packages also do `import *` from files which polutes the package namespace with all kinds of stuff you really don't want in there. For example, the popular Requests package: >>> import requests >>> requests.logging <module 'logging' from '/usr/lib/python2.7/logging/__init__.pyc'> The logging module is not a public part of requests' API. It's just there because requests uses it internally. So to answer your question, I'd say it's just common practice. If it's undocumented in Python, you should pretend it doesn't exist.
- Grue3 11y agoPython's re has nothing on CL-PPCRE [1] though. The ability to build up a "regular expression" from S-expressions is just too useful. [1] http://weitz.de/cl-ppcre/ http://weitz.de/cl-ppcre/
- nanny 11y agoIt was also twice as fast as Perl in benchmarks at one point or another.
- kbenson 11y agoThat's not exactly hard to do, depending on features. There's a definite trade-off between features and the type of regex engine that can be implemented.
- DasIch 11y agoIt seems to me that one should be able to detect which features are actually used in a regular expression and choose one of multiple different underlying implementations based on that.
- kbenson 11y agoNewer versions of Perl actually support a pluggable regex engine system, so you can use specific regex engines for specific tasks.
- Twirrim 11y agoIf you hook in libpcre direct it actually comes with its own JIT these days. http://sljit.sourceforge.net/pcre.html http://sljit.sourceforge.net/pcre.html
- draven 11y agoHaving the compiler available at runtime kinda helps too!
- junke 11y ago
- stefantalpalaru 11y ago> it's one of the best of all dynamic languages I would argue The author should learn about PCRE. I wrote a Python wrapper for it that includes a drop-in 're' substitute: https://github.com/stefantalpalaru/morelia-pcre https://github.com/stefantalpalaru/morelia-pcre
- larkinrichards 11y agoI believe there is a small bug in the final example, in the tokenize definition it references 'self' where I believe it should reference 'scanner'.
- deleted 11y ago[deleted]
- nichochar 11y agoAmazing! This Scanner object seems so beautifully pythonic and intuitive. It's a shame it's not documented honestly: it would be a great way for beginners in python to write language parsers