4 ms·
> "one code point in unicode does not necessarily map to one character on the screen." Also, importantly, some characters are not represented at all. For examp
by nanis 5y ago
> "one code point in unicode does not necessarily map to one character on the screen."
Also, importantly, some characters are not represented at all. For example, there is no codepoint or combination of characters that distinguishes a capital Turkish dotless I from its identical looking but conceptually different sibling the Latin capital I. Similarly for capital Turkish dotted İ.
I find it extremely weird that two codepoints couldn't have been spared, yet we have gradations of skin tone in emoji. One can use composition to deal with the round-trip from i -> İ -> i (within a closed ecosystem), but even that fails when it comes to ı -> I -> ı.
A case in point are the product pages for my friend's book on Amazon. Compare "Sınır Ötesi" to "Sinir Ötesi". The former means "beyond borders" whereas the latter means "exceedingly irritating".
Yet, because there is no unambiguous representation of the Turkish I's, the rendering makes an assumption on the basis of the domain. Note that all other non-US ASCII letters involved are rendered correctly.
\c[PERSON FROWNING, ZERO WIDTH JOINER, PROGRAMMER]
[tr]: https://www.amazon.com.tr/Gezging%C3%B6z-S%C4%B1n%C4%B1r-T%C3%BCrkiye-Miras%C4%B1-Rehberi/dp/6257631181/ https://www.amazon.com.tr/Gezging%C3%B6z-S%C4%B1n%C4%B1r-T%C...
[us]: https://www.amazon.com/Gezging%C3%B6z-Sinir-T%C3%BCrkiye-Mirasi-Rehberi/dp/6257631181 https://www.amazon.com/Gezging%C3%B6z-Sinir-T%C3%BCrkiye-Mir...
- samatman 5y agoThis is an outcome of the historic process which gave us modern Unicode. The Turkish alphabet ISO 8859-9 was published in 1988, and it doesn't distinguish between I and I, since why would it? It's a byte encoding designed specifically to accommodate Turkish. Unicode has a principle of adopting whatever makes it easy to convert from the preferred legacy encoding to Unicode, and for Turkish that meant using the same code point for I in both Turkish and practically everywhere else, since I is in ASCII and the bulk of Latin-based character encodings are based on ASCII. I point this out because it could indeed have been the case that they ended up with separate encodings, which would have solved your problem. But it wouldn't have solved the actual problem, which is the collation and casing are locale-specific according to Unicode, and indeed they must be. Example: in Dutch the letters 'ij' are one letter for collation purposes, as well as for casing, see https://en.wikipedia.org/wiki/IJsselstein https://en.wikipedia.org/wiki/IJsselstein for an example. This is of course impossible to get correct if the application doesn't know it's dealing with Dutch, which it often can't know. The very same problem you have with Turkish.
- schiffern 5y ago>The very same problem you have with Turkish. Not really the same. The brokenness of more complex text transformations (collation, alphabetization, upper-casing, lower-casing) seems a bit less catastrophic than breaking the fundamental ability to unambiguously draw the character. If you don't think that's an important goal.... then what are "character sets" for anyway?
- bombcar 5y agoThat brings up the philosophical problem - is the English word taco the same as the Spanish word taco and if they aren’t the same, are the characters the same, or should there be an a for every language?
- samatman 5y agoWhich semantic HTML at least attempts to solve with the `span lang` construct, although "taco" is maybe not the ideal example, but saying e.g. <span lang=fr>Jean</span> could in principle tell a screen reader to pronounce this name correctly, rather than the equally-correct English pronunciation which is a homophone of "gene". This is another question which doesn't belong in the Unicode character specification, but should be resolved when necessary at a higher level of abstraction.
- bombcar 5y agoExactly. Unicode’s “job” is to provide a code point for every letter - which necessary involves some trade offs such as whether a character is duplicated or not depending on the history of the region (which depends on all sorts of historical accidents such as whether keyboards were even available). People shouldn’t expect to know what language a Unicode string is in without external help, even if heuristics work often.
- nanis 5y ago> Exactly. Unicode’s “job” is to provide a code point for every letter And in this case Unicode fails catastrophically by not providing codepoints for the uppercase of "ı" and the lowercase of "İ" while potentially allowing for ethnically sensitive renderings of shit.