5 ms·
Sure, but that's up to whoever is writing the regular expression. The standard '.' (match any character) in a regexp matches an Unicode code point, not a graph
by moefh 4y ago
Sure, but that's up to whoever is writing the regular expression.
The standard '.' (match any character) in a regexp matches an Unicode code point, not a grapheme cluster. To match a grapheme cluster, you have to use "\X", which is not universally supported.
For example, in Python 3, the builtin module module "re" doesn't support "\X", you have to install the "regex" module for it:
# text is 'e' followed by U+0301 (combining acute accent):
text = b'e\xcc\x81'.decode('utf8')
print(f'text: "{text}"')
import re
print(re.match('^(.)(.)$', text).groups()) # prints "('e', '´')"
import regex # must be installed
print(regex.match('^(\X)$', text).groups()) # prints "('é',)"
- bmn__ 4y ago> To match a grapheme cluster, you have to use "\X", which is not universally supported. You have the words "supported" and "implemented" mixed up. Kernighan claims Unicode support, so he is required by the standard the implement \X. If a software does not implement \X, then it is not compliant, and it would be very wrong to say it supports Unicode. Does anyone have a deeplink showing the evidence for awk?
- moefh 4y ago> so he is required by the standard the implement \X. Which standard is that? If you're talking about Unicode, the "standard" for regular expressions[1] is an "Unicode Technical Standard", which according to itself isn't required for Unicode conformance: > A Unicode Technical Standard (UTS) is an independent specification. Conformance to the Unicode Standard does not imply conformance to any UTS. So awk can claim Unicode support without supporting "\X" (like many regex engines). If you're talking about POSIX, its regex chapter[2] doesn't mention "\X". In any case I don't think awk claims to conform to POSIX. [1] https://unicode.org/reports/tr18/ https://unicode.org/reports/tr18/ [2] https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1_chap09.html https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
- burntsushi 4y agoThis is just false. UTS#18 specifies multiple levels of Unicode support. \X is part of level 2. It is perfectly valid to generally say "has Unicode support" even if it's just Level 1, assuming you document somewhere what precisely is supported. For example, I regularly say that Rust's regex crate has Unicode support. But it does not support \X. It's more precisely documented here: https://github.com/rust-lang/regex/blob/master/UNICODE.md https://github.com/rust-lang/regex/blob/master/UNICODE.md
- deleted 4y ago[deleted]