4 ms·
This particular case seems odd to me because INFO is an English word, and ınfo is not.
by TwoBit 6y ago
This particular case seems odd to me because INFO is an English word, and ınfo is not.
- wongarsu 6y agoYou could make a case that Unicode should have different "i" characters for different languages. Then you could do all transformations unambiguously. On the other hand almost everyone abuses the minus sign as a dash, and treats the apostrophe and the prime sign (signifying feet or minutes) as interchangeable, so in all likelihood they would constantly use the wrong i too.
- heavenlyblue 6y agoPretty sure that’s not true. When you switch your keyboard you will have a proper i character in another language unless your keymap is broken. How do you think Chinese, Russians or Greek type their characters?
- tzot 6y agoThe grandparent obviously meant “latin i”; none of the three languages you mention have any latin letters, but at least Russian and Greek have some lowercase and some more uppercase letters with the same glyph/shape as latin ones.
- heavenlyblue 6y agoYeah, and those similar glyphs are not available on their own language keyboard.
- wongarsu 6y agoI frequently type German with a US layout with dead keys (so I can type "a to get ä). I also imagine that most Turkish developers type English on a Turkish layout, since Turkish contains all characters used by English.
- anticensor 6y agoI have a better solution: use combining characters COMBINING DOT ABOVE (which already exists) and DELETE DOT ABOVE (which needs to be added into Unicode), which would manipulate "I" into "İ" and "i" into "ı" respectively. Those combining characters would also work perfectly with j too.
- estebank 6y agoThe only issue I can see is with people working in a Turkish locale writing Latin text producing, let's say English blogposts with the wrong i and I. I still think that this should have been done this way though...
- anticensor 6y agoIndeed. LATIN SMALL LETTER I + DELETE DOT ABOVE becomes LATIN CAPITAL LETTER I + DELETE DOT ABOVE in uppercase, which then becomes LATIN SMALL LETTER I + DELETE DOT ABOVE back in lowercase. The same thing applies to LATIN CAPITAL LETTER I + COMBINING DOT ABOVE. Survives infinite number of case conversions.
- johnwalkr 6y agoWell a round-trip or two could still be ambiguous which could easily fail when comparing strings later in some edge case. Especially when we can't even consistently agree to use by-application, by-OS, by-language and by-locale settings consistently. I don't have a solution, just pointing out that this is a really challenging problem to fully solve.
- kps 6y ago> On the other hand almost everyone abuses the minus sign as a dash Unicode calls it HYPHEN-MINUS. It does also have an unambiguous ‘−’ MINUS SIGN as well as ‘‐’ U+2010 HYPHEN and the various dashes, but most people use bad keyboard layouts.
- josefx 6y ago> You could make a case that Unicode should have different "i" characters for different languages. And different "SS" for any case where the lowercase was an sz, of course at some point Germany introduced an uppercase SZ character to avoid that round trip loss issue, but we still have tons of text that use the old sz -> SS conversion. Also note that "y" in Germany, not all German speaking countries follow the same rules for sz, some dropped it entirely. We basically need something like the time zone database to have even a snowballs chance in hell to handle text correctly.