3 ms·
You could do it the way Raku does. It's implementation defined. (Rakudo on MoarVM) The way MoarVM does it is that it does NFG, which is sort of like NFC except
by b2gills 6y ago
You could do it the way Raku does.
It's implementation defined.
(Rakudo on MoarVM)
The way MoarVM does it is that it does NFG, which is sort of like NFC except that it stores grapheme clusters as if they were negative codepoints.
If a string is ASCII it uses an 8bit storage format, otherwise it uses a 32bit one.
It also creates a tree of immutable string objects.
If you do a substring operation it creates a substring object that points at an existing string object.
If you combine two strings it creates a string concatenation object. Which is useful for combining an 8bit string with a 32bit one.
All of that is completely opaque at the Raku level of course.
my $str = "\c[FACE PALM, EMOJI MODIFIER FITZPATRICK TYPE-3, ZWJ, MALE SIGN, VARIATION SELECTOR-16]";
say $str.chars; # 1
say $str.codes; # 5
say $str.encode('utf16').elems; # 7
say $str.encode('utf16').bytes; # 14
say $str.encode.elems; # 17
say $str.encode.bytes; # 17
say $str.codes * 4; # 20
#(utf32 encode/decode isn't implemented in MoarVM yet)
.say for $str.uninames;
# FACE PALM
# EMOJI MODIFIER FITZPATRICK TYPE-3
# ZERO WIDTH JOINER
# MALE SIGN
# VARIATION SELECTOR-16
The reason we have utf8-c8 encode/decode is because filenames, usernames, and passwords are not actually Unicode.
(I have 4 files all named rèsumè in the same folder on my computer.)
utf8-c8 uses the same synthetic codepoint system as grapheme clusters.