3 ms·
And yet there is still no unambiguous lower case "I" or upper case "i".
by nanis 4y ago
And yet there is still no unambiguous lower case "I" or upper case "i".
- ClumsyPilot 4y agothats the job of a font, not encoding
- nanis 4y agoNo, the fact that there is no codepoint that makes those mappings ambiguous is due to the way Unicode decided to save to codepoints for seemingly no good reason. What should be _the_ value of `"I".lower()`? Or, "i".upper()? And please don't bring up locales. The whole point of accepting the complexity of Unicode is to be able to take a document which stands on its own without external references. > Early character encodings also conflicted with one another. That is, two encodings could use the same number for two different characters, or use different numbers for the same character. > The Unicode Standard provides a unique number for every character, no matter what platform, device, application or language.[1] Those statements are outright lies: Unicode does not provide a unique number fpr "upper case Turkish dotless i". Nor does it provide one for "lower case Turkish dotted i". If it did, it would be possible to correctly map "i" to "I" or "İ" and "I" to "i" or "ı" without having to know anything other than the source codepoint. The font does not even come into play here. [1]: https://unicode.org/standard/WhatIsUnicode.html https://unicode.org/standard/WhatIsUnicode.html
- deleted 4y ago[deleted]
- Kwpolska 4y agoUnicode is for representing text, not allowing arbitrary manipulation of it. It isn’t the job of Unicode to encode those relationships. Also, the Turkish `i` stuff is just the tip of the iceberg. Should Unicode be able to round-trip `'ß'.upper().lower()`? Keeping the existing capitalization of ß → SS, you need to define a "uppercase S that used to be ß" character. Then there’s the Dutch `ij`, in which both characters are either uppercase or lowercase (`Ij` at the start of a word is incorrect). There’s a ligature in Unicode, but it’s only for compatibility with some legacy keymaps. But is there a point in adding a new version of "S" that a lot of software would not recognize as equivalent to the plain old ASCII "S" (and one might end up far away from a ß due to copy-pasting or stuff), bringing weird bugs and security issues? Should the Dutch throw out all their keyboards just so they get a new key for the special IJ ligature?
- msla 4y ago> Keeping the existing capitalization of ß → SS, you need to define a "uppercase S that used to be ß" character. Even more complicated, there is a capital form of ß: https://en.wikipedia.org/wiki/%C3%9F https://en.wikipedia.org/wiki/%C3%9F > Until 2017, there was no official capital form of ⟨ß⟩; a capital form was nevertheless frequently used in advertising and government bureaucratic documents.[10]: 211 In June of that year, the Council for German Orthography officially adopted a rule that ⟨ẞ⟩ would be an option for capitalizing ⟨ß⟩ besides the previous capitalization as ⟨SS⟩ (i.e., variants STRASSE and STRAẞE would be accepted as equally valid).[11] [12] Prior to this time, it was recommended to render ⟨ß⟩ as ⟨SS⟩ in allcaps except when there was ambiguity, in which case it should be rendered as ⟨SZ⟩. The common example for such a case was IN MASZEN (in Maßen "in moderate amounts") vs. IN MASSEN (in Massen "in massive amounts"), where the difference between the spelling in ⟨ß⟩ vs. ⟨ss⟩ could actually reverse the conveyed meaning.[citation needed] No character encoding standard can save you from that kind of complexity. You need special-case code on a language-by-language basis.
- Kwpolska 4y agoThe capital ẞ is a recent invention, it’s in Unicode since 2008. While it would solve the issue, most programming languages and other tools will convert it to SS (though things are changing; Gboard now offers an actual uppercase ẞ instead of the SS it had some years ago). I don’t know what actual humans think, but I suppose the no-capital-ß-exists position is ingrained in many German speakers.
- msla 4y agoI wasn't saying it was a solution, I was saying it was a further complication. Also, Unicode didn't invent it. Germans invented it.
- jrochkind1 4y ago> The whole point of accepting the complexity of Unicode is to be able to take a document which stands on its own without external references. What makes you think this is the whole point? I don't believe that "whole point" is actually possible with global human languagues as actually used, nor do I think those behind unicode historically or presently have considered this the "whole point". Do you have any reference suggesting this was meant to be the "whole point" of unicode? I think it's actually a pretty amazing success that unicode does give us algorithms for manipulating text in various ways, that actually work pretty darn well... but yes, they sometimes require locale parameters.
- mcswell 4y ago"Those statements are outright lies: Unicode does not provide a unique number fpr "upper case Turkish dotless i". Nor does it provide one for "lower case Turkish dotted i"." By design, and I would argue that this is the correct design. What it does provide is a unique code point for "upper case I", "upper case dotted I", "lower case dotted i" and "lower case dotless i".
- nanis 4y agoYet, for example, it distinguishes upper and lower case versions iota[1] from "upper case i" and "lower case dotless i". A similar provision could have been made for upper and lower case versions of "Turkish dotless i" and "Turkish dotted i". But it wasn't. That means "ι".upper().lower() produces "ι" but "ı".upper().lower() produces "i" (unless you do it on a computer set to a Turkish locale: >>> "ι".upper().lower() 'ι' >>> "ı".upper().lower() 'i' Note we have: LATIN CAPITAL LETTER IOTA LATIN SMALL LETTER IOTA GREEK CAPITAL LETTER IOTA GREEK SMALL LETTER IOTA MODIFIER LETTER SMALL IOTA TURNED GREEK SMALL LETTER IOTA CYRILLIC CAPITAL LETTER IOTA CYRILLIC SMALL LETTER IOTA APL FUNCTIONAL SYMBOL IOTA MATHEMATICAL CAPITAL IOTA MATHEMATICAL BOLD SMALL IOTA MATHEMATICAL ITALIC CAPITAL IOTA MATHEMATICAL ITALIC SMALL IOTA MATHEMATICAL BOLD ITALIC CAPITAL IOTA MATHEMATICAL BOLD ITALIC SMALL IOTA MATHEMATICAL SANS-SERIF BOLD CAPITAL IOTA MATHEMATICAL SANS-SERIF BOLD ITALIC CAPITAL IOTA MATHEMATICAL SANS-SERIF BOLD ITALIC SMALL IOTA That's lotta iotas[1]. There's more[2]: GREEK SMALL LETTER IOTA WITH DASIA GREEK SMALL LETTER IOTA WITH PSILI AND VARIA GREEK SMALL LETTER IOTA WITH DASIA AND VARIA ... The point is that almost all of these variations on the theme with explicit duplicates get their upper and lower case distinct codepoints, yet "lower case Turkish dotted i" needs to for some reason map to 0x69 in ASCII and "upper case Turkish dotless i" needs to map to 0x49 ASCII. It is stupid and directly contradicts the statement: > The Unicode Standard provides a unique number for every character, no matter what platform, device, application or language.[3] The statement on the Unicode.org web site cannot be attributed to ignorance. [1]: https://en.wikipedia.org/wiki/Iota https://en.wikipedia.org/wiki/Iota [2]: https://www.unicode.org/charts/PDF/U1F00.pdf https://www.unicode.org/charts/PDF/U1F00.pdf [3]: https://unicode.org/standard/WhatIsUnicode.html https://unicode.org/standard/WhatIsUnicode.html
- cryptonector 4y agoThat can't be fixed. You just have to know if you're in a locale that wants one or another rule for that. Well, I suppose the UC could still introduce new codepoints for I and i that specifically for Turkish, but it wouldn't be convenient for users of Turkish.