4 ms·
Additionally there's a newer module also named 'regex': https://pypi.python.org/pypi/regex https://pypi.python.org/pypi/regex
by andreasvc 11y ago
Additionally there's a newer module also named 'regex': https://pypi.python.org/pypi/regex https://pypi.python.org/pypi/regex
- rspeer 11y agoAnd this newer 'regex' is actually really good at tricky cases such as matching word boundaries. (An apostrophe or a non-ASCII character is not necessarily a word boundary!)
- andreasvc 11y ago> An apostrophe or a non-ASCII character is not necessarily a word boundary! I don't see how a regular expression library could help with that (other than proper Unicode support), because word boundaries are a language-specific, linguistic problem; i.e., you will need to supply a list of possible contractions anyway. Tokenization of natural language text may appear like a straightforward and solved problem, but there are actually lots of messy details to get right.