3 ms·
UTF-32 is a fixed-length encoding of Unicode[1], so it does simplify things a lot for a regex engine. [1] At least when talking about code points, which is wha
by moefh 4y ago
UTF-32 is a fixed-length encoding of Unicode[1], so it does simplify things a lot for a regex engine.
[1] At least when talking about code points, which is what matters for regular expressions (unless you want stuff like \X with is not universally supported).
- kqr 4y agoWait, does Unicode not have multiple representations even of simple things like the letter "ä"? Then you definitely need to handle actual characters/glyphs in regexes.
- moefh 4y agoSure, but that's up to whoever is writing the regular expression. The standard '.' (match any character) in a regexp matches an Unicode code point, not a grapheme cluster. To match a grapheme cluster, you have to use "\X", which is not universally supported. For example, in Python 3, the builtin module module "re" doesn't support "\X", you have to install the "regex" module for it: # text is 'e' followed by U+0301 (combining acute accent): text = b'e\xcc\x81'.decode('utf8') print(f'text: "{text}"') import re print(re.match('^(.)(.)$', text).groups()) # prints "('e', '´')" import regex # must be installed print(regex.match('^(\X)$', text).groups()) # prints "('é',)"
- bmn__ 4y ago> To match a grapheme cluster, you have to use "\X", which is not universally supported. You have the words "supported" and "implemented" mixed up. Kernighan claims Unicode support, so he is required by the standard the implement \X. If a software does not implement \X, then it is not compliant, and it would be very wrong to say it supports Unicode. Does anyone have a deeplink showing the evidence for awk?
- moefh 4y ago> so he is required by the standard the implement \X. Which standard is that? If you're talking about Unicode, the "standard" for regular expressions[1] is an "Unicode Technical Standard", which according to itself isn't required for Unicode conformance: > A Unicode Technical Standard (UTS) is an independent specification. Conformance to the Unicode Standard does not imply conformance to any UTS. So awk can claim Unicode support without supporting "\X" (like many regex engines). If you're talking about POSIX, its regex chapter[2] doesn't mention "\X". In any case I don't think awk claims to conform to POSIX. [1] https://unicode.org/reports/tr18/ https://unicode.org/reports/tr18/ [2] https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1_chap09.html https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
- burntsushi 4y agoThis is just false. UTS#18 specifies multiple levels of Unicode support. \X is part of level 2. It is perfectly valid to generally say "has Unicode support" even if it's just Level 1, assuming you document somewhere what precisely is supported. For example, I regularly say that Rust's regex crate has Unicode support. But it does not support \X. It's more precisely documented here: https://github.com/rust-lang/regex/blob/master/UNICODE.md https://github.com/rust-lang/regex/blob/master/UNICODE.md
- deleted 4y ago[deleted]
- Pelam 4y agoAt the end of this comment there is what I believe is a single grapheme cluster. On disk this single "letter" occupies 73 bytes. Surprisingly large number of tools and editors know how to work with things like these and render them at least somehow I think I once created one that was about a kilobyte. Is there an upper limit? I created it using this page https://glitchtextgenerator.com/ https://glitchtextgenerator.com/ The 73 byte X: x̧̡̬̘͓̖̲̻̻̲̠̪̻͓͙̜̂̓̊̔̀̀͗̑̀̅̀̂̚͘̕̚͘͢͜͠