3 ms·
Aiui the Unicode problem is even more complex than most folk realize for the reasons I explain in this comment. ----------------------------------------- On t
by raiph 11y ago
Aiui the Unicode problem is even more complex than most folk realize for the reasons I explain in this comment.
-----------------------------------------
On the final technical destination implied by Unicode:
The final technical destination implied by the Unicode standard is correct handling of "grapheme clusters".[1]
-----------------------------------------
On the half way point at which we've arrived:
In the 80s, there were a bazillion incompatible "character" sets and encodings of those "character" sets. Unicode was a response to this.
The half way point envisaged by the original Unicode design was to firmly establish a single overall "character set" view with a handful of encodings.
Imo we've arrived at this half way point.
-----------------------------------------
On the string types of languages:
Programming languages designed before Unicode emerged adopted the "byte=character" view to determine the character unit of their default string type, leaving higher level views to library code and functions. This is the level Python began at. But the more established Unicode becomes, the more important the higher level views become.
Newer languages have mostly been designed to optimize for Unicode's half way point (but not the eventual destination). They typically support the byte view in some form but adopt the "codepoint=character" view to determine the character unit of their default string type. Aiui, this is where Python 3 is at.
A handful of newer languages continue to support the byte and codepoint views in some form but have adopted the highest level "grapheme=character" view for the character unit of their default string type. I think the most visible new language that's been forward-thinking enough to adopt the final destination view (grapheme=character) is Swift.
-----------------------------------------
[1] Folks trying to understand this point should find it helpful to start with the deceptively brief hint inherent in http://www.unicode.org/glossary/#Grapheme http://www.unicode.org/glossary/#Grapheme which defines "what a user thinks of as a character".
Next, I'll quote the perluniintro[2] document:
> A Unicode logical "character" can actually consist of more than one internal actual "character" or code point. For Western languages, this is adequately modelled by a base character (like LATIN CAPITAL LETTER A ) followed by one or more modifiers (like COMBINING ACUTE ACCENT ). This sequence of base character and modifiers is called a combining character sequence. Some non-western languages require more complicated models, so Unicode created the grapheme cluster concept, which was later further refined into the extended grapheme cluster.
[2] http://perldoc.perl.org/perluniintro.html http://perldoc.perl.org/perluniintro.html