3 ms·
> I'm getting tired of this argument Perhaps this is because Unicode is an exhaustive standard. :) Seriously, I hear ya. > every argument had the same hole t
by raiph 11y ago
> I'm getting tired of this argument
Perhaps this is because Unicode is an exhaustive standard. :)
Seriously, I hear ya.
> every argument had the same hole that I see in your post: What, exactly, is a "character"? Please answer me that
I'd say the answer varies depending on who's talking, what they're talking about, and who's listening.
If it's a programmer interested in listening to the Unicode consortium and interested in what Unicode.org documents define for "what a user thinks of as a character" then the answer is, according to those documents, a "grapheme".
> then you can write a proposal for how the API should work, exactly
Right.
It can get, and has indeed gotten, better than that; one can write specifications, build reference implementations, and try things out in battle for a few years. Several text processing systems (including some programming languages) have gone this route and their results should be taken in to account. (ICU is perhaps the go to reference implementation.)
> and why this works with Indic and German and Korean and everything else.
For designs and implementations of text processing systems that claim an aspiration of progressing toward fully following the Unicode specification, the simplest answer to "why does it work?" (or "why it is expected to work") with Indic and German and Korean and everything else is of course the Unicode standard itself.
For those of us who aren't blessed with the patience of a saint and thus haven't read the entire spec from start to finish a few times, one has to rely to a degree on those who do, even if it's really hard to follow what the heck they're saying.
> You're trying to prevent people from slicing up grapheme clusters, which is a noble goal.
Given that grapheme clusters correspond to what the Unicode standard specifies for identifying and preserving the integrity of "what a user thinks of as a character", it seems less a noble goal than a fundamental eventual goal for any text processing system that explicitly claims to aspire to support the world's human languages via Unicode compliance.
> we don't end up slicing grapheme clusters very often in practice.
Which "we" are you referring to?
The western world has dominated text processing to date does. This must not mean that "we", by which I mean all humans, would best continue practices that marginalize those with native languages not well served. (By ignoring the central importance of "what a user thinks of as a character"; again, I'm quoting Unicode consortium documentation.)
Or, more parochially, what about the "we" that's just programmers who want to be able to increase their confidence that they are properly processing Unicode compliant text?
> In the meantime, formats like JSON and XML are defined in terms of code points ...
Code points are the middle level of Unicode, sandwiched between byte encodings at the bottom level and graphemes at the top. There's nothing wrong with a text processing system stopping at the middle level provided folk don't lose sight of the fact that it stops at the middle level.
Thinking that code points are "what a user thinks of as a character" mostly works out OK when processing text that's English or the like, but is definitively not Unicode compliant to the degree it's used too broadly.
> Can you tell me what the correct behavior is?
I would suggest that the most appropriate authority is the Unicode consortium and its Unicode standard.
> Might you end up with bugs in programs because len(x + y) != len(x) + len(y)?
Oh for sure. But it's part of the Unicode standard. It's one of several apparently strange discontinuities that the Unicode effort introduced last century in an attempt to deal with the combination of human language reality and computer system realities. They may have chosen the wrong sweetspot. But now Unicode is a significant part of reality. I, for one, don't think it's wise or even realistic to abandon the Unicode standard as it is...