5 ms·
If you want to work with these in Python, use the "regex" package. [1] Not to be confused with the "re" module in the standard library. Also, maybe there shoul
by rspeer 8y ago
If you want to work with these in Python, use the "regex" package. [1] Not to be confused with the "re" module in the standard library.
Also, maybe there should be a (2016) or even a (2004) in the title! (2016 is when this document was last revised, and 2004 is when it became a Unicode standard.)
[1] https://pypi.org/project/regex/ https://pypi.org/project/regex/
- xenomachina 8y agoWasn't "regex" also the name of one of re's predecessors?
- jwilk 8y agoYes. It was deprecated in 1.5(!) and finally removed in 2.5: https://docs.python.org/release/1.5.1/lib/module-regex.html https://docs.python.org/release/1.5.1/lib/module-regex.html https://docs.python.org/2/whatsnew/2.5.html#new-improved-and-removed-modules https://docs.python.org/2/whatsnew/2.5.html#new-improved-and...
- jwilk 8y agoI don't think the regex package implements this. At least the hex notation is not supported: >>> import regex >>> regex.compile(r'\u{3040}') Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/usr/lib/python3/dist-packages/regex.py", line 345, in compile return _compile(pattern, flags, kwargs) File "/usr/lib/python3/dist-packages/regex.py", line 507, in _compile caught_exception.pos) _regex_core.error: incomplete escape \u at position 2
- rspeer 8y agoHmm, okay. It seems to just use Python's unicode escape syntax, instead of what the standard says. But the package does support things like the complex character classes, and identifying word boundaries. (The algorithm for identifying word boundaries in most languages without ASCII/English assumptions is quite complex and useful, as I can say from having tried to reinvent half of it before learning about the standard and the regex package. It's not a panacea -- it won't do anything useful with languages where word boundaries require lexical knowledge, like Chinese, Japanese, and Thai -- but other than that it handles all the edge cases you never would have thought of.)
- burntsushi 8y agoYeah, basically, the standard is a little sneaky on this point. It doesn't actually require a specific syntax, but rather, simply that being able to write Unicode codepoints in hexadecimal representation is possible. Notice that it says, "... shall supply a mechanism ..." rather than "this mechanism." Of course, it's probably a good idea to follow the sample syntax provided. :-)