3 ms·
> An apostrophe or a non-ASCII character is not necessarily a word boundary! I don't see how a regular expression library could help with that (other than prop
by andreasvc 11y ago
> An apostrophe or a non-ASCII character is not necessarily a word boundary!
I don't see how a regular expression library could help with that (other than proper Unicode support), because word boundaries are a language-specific, linguistic problem; i.e., you will need to supply a list of possible contractions anyway.
Tokenization of natural language text may appear like a straightforward and solved problem, but there are actually lots of messy details to get right.