3 ms·
I agree that for CJK-natives it's not such a big deal probably (unless they live their life in more than one of those languages). For people like me who primari
by cehrlich 5y ago
I agree that for CJK-natives it's not such a big deal probably (unless they live their life in more than one of those languages). For people like me who primarily use their computer in English but also do some stuff in Japanese every now and then it's very frustrating. Of course at this point I know what's going on when 直 or whatever looks wrong, but it's still frustrating.
OSs and Browsers having their own logic for it actually makes things _worse_ in some cases. Windows is especially bad (different types of UI elements care about different settings or don't care at all, so good luck having apps render correctly if you don't change your entire OS locale), and Chrome is pretty bad too, again especially on Windows. Overall MacOS/iOS and Safari does the best job by far.
The failed attempt at Han Unification[1] is the worst decision the Unicode people have ever made.
[1] https://en.wikipedia.org/wiki/Han_unification https://en.wikipedia.org/wiki/Han_unification
- jandrese 5y agoI disagree. UCS-2 was the worst decision the Unicode consortium ever made. Han unification is merely a side effect of this bad decision.
- nanis 5y ago> The failed attempt at Han Unification[1] is the worst decision the Unicode people have ever made. At first I nodded my head in agreement, but then I decided I still think the failure to include separate code points for "lower case Turkish dotted I" and "upper case Turkish dotless I" is worse. You can't have 'ı' ≡ lc( uc 'ı' ) unless you already know you are processing Turkish ... completely unnecessary complication.
- naniwaduni 5y agoTurkish I unification, at least, wasn't a decision the Unicode people made, they inherited the mistake from earlier encodings. Given that those already existed, the alternative to having broken casefolding was, essentially, break all mixed Turkish documents transcoded from cp857 containing both "I" and "i" in non-Turkish functional directives, i.e. you'd necessarily break things like HTML documents without consistent tag casing.
- nanis 5y agoI am having a hard time seeing how having the option of distinct codepoints would break anything. Consider, İ/i where it is _possible_ to do lossless case conversion: > lc( 'İ' ) becomes i followed by COMBINING DOT ABOVE which means uc(lc 'İ') becomes LATIN CAPITAL LETTER I WITH DOT ABOVE as a by product of the fact that perl6 deals in graphemes[1] say 'İ' eq 'İ'.lc.uc.lc.uc; True If an extra codepoints existed for Turkish dotted I, such contortions would not be necessary and this would have had no implications for existing working code at the time (nothing says those codepoints must be used, they just give smart software options). Now, there is nothing one can do with I/ı that will make 'ı' eq 'ı'.uc.lc.uc.lc true without extra information. If codepoints existed, then such special casing and carrying around extra information would not have been necessary. Also note: # The letter Ö is not considered to be a variant of the letter O, # and is a separate letter in the Swedish alphabet. The former # character is, however, the accepted alternative in contexts where # Ö cannot be used. Earlier practice substituted OE, which is no # longer recommended but will still be encountered. # U+00F6 # LATIN SMALL LETTER O WITH DIAERESIS It should not have been too hard to say "The letter İ is not considered to be a variant of the letter I" and vice versa for the lower case versions. [1]: https://www.nu42.com/2017/02/for-your-eyes-only.html https://www.nu42.com/2017/02/for-your-eyes-only.html
- cehrlich 5y agocall it a tie