41 ms·
In the Turkish locale, "INFO".lower() != "info"
- anticensor 6y agoCorrect: you would get "ınfo", "warnıng" and "crıtıcal" in Turkish and in Azerbaijani.
- mapgrep 6y agoFurther context: https://en.m.wikipedia.org/wiki/Dotted_and_dotless_I https://en.m.wikipedia.org/wiki/Dotted_and_dotless_I Did not know Istanbul is actually İstanbul.
- gvx 6y agoMe neither. I did know it's not Constantinople, though.
- anticensor 6y agoConstantinople (Fatih) is the capital town of Eistipolis (Istanbul).
- tantalor 6y agohttps://bugs.python.org/issue1524081 https://bugs.python.org/issue1524081 > KeyError: 'Info'
- scrollaway 6y agoIve long thought programming languages need a "localizable string" (Aka user-facing string) type, different from regular utf8 strings. Something like what gettext and other i18n libraries fake for you, but native to the language. Behaviour like this is definitely a good reason why: sorting, changing case, etc should be consistent when dealing with strings used as constants and identifiers, but Python's .lower() behaviour makes sense in a localizable string context.
- DougBTX 6y agoAlong the lines of this? https://docs.microsoft.com/en-us/dotnet/api/system.globalization.cultureinfo.invariantculture?view=netcore-3.1 https://docs.microsoft.com/en-us/dotnet/api/system.globaliza...
- wongarsu 6y ago.NET is one of the few ecoecosystems to get this right. It offers the invariant culture for identifier-like things, "fr" for French language and "fr-FR" for French language in France, allowing you to specify your intention to every string-modifying function. Support at the type level would be a lot less verbose, but support at the function level is already much better than many other popular languages.
- kanox 6y agoIt would be great if strings and especially date-time values always carried locale and timezone information with them. It would take slightly more memory but not significant on modern machines.
- wongarsu 6y agoPutting the locale information on the string sounds like a good idea. However I'm not sure how that should handle combined strings with components from different locales. For example `logLevel + ": " + logMessage` might produce "info: bağlantı kesildi" in Turkish. How to annotate that? Neither English nor Turkish would work correctly, each would produce the wrong result when uppercasing. You could treat it as a series of string slices with different locales `[("info", "en"), (": ", ""), ("bağlantı kesildi", "tr")]`. That would work correctly, and you could now uppercase each slice according to its appropriate locale, but it wouldn't really be low overhead anymore. Maybe still worth it. It would be an interesting approach that might even be able to be implemented pretty seamlessly as a library in some languages (C++ or rust for example)
- layer8 6y ago
- chippy 6y agohttps://garygregory.wordpress.com/2015/11/03/java-lowercase-conversion-turkey/ https://garygregory.wordpress.com/2015/11/03/java-lowercase-... In the Turkish locale, the Unicode LATIN CAPITAL LETTER I becomes a LATIN SMALL LETTER DOTLESS I. That’s not a lowercase “i”.
- deleted 6y ago[deleted]
- geofft 6y agoIn C (POSIX.1-2008, specifically), there's tolower_l() and the rest of the _l functions for this use case, which take a locale as an argument. That let's you ask for the English (or even "C locale") lowercase versions of these English words, even when your process's current locale is Turkish. https://www.man7.org/linux/man-pages/man3/tolower_l.3.html https://www.man7.org/linux/man-pages/man3/tolower_l.3.html
- adamjb 6y agoThe mention of _l functions reminded me of this gloriously over the top git message/rant. "Those not comfortable with toxic language should pretend this is a religious text." https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f027338b0fab0f5078971fbe https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]
- TwoBit 6y agoThis particular case seems odd to me because INFO is an English word, and ınfo is not.
- wongarsu 6y agoYou could make a case that Unicode should have different "i" characters for different languages. Then you could do all transformations unambiguously. On the other hand almost everyone abuses the minus sign as a dash, and treats the apostrophe and the prime sign (signifying feet or minutes) as interchangeable, so in all likelihood they would constantly use the wrong i too.
- heavenlyblue 6y agoPretty sure that’s not true. When you switch your keyboard you will have a proper i character in another language unless your keymap is broken. How do you think Chinese, Russians or Greek type their characters?
- tzot 6y agoThe grandparent obviously meant “latin i”; none of the three languages you mention have any latin letters, but at least Russian and Greek have some lowercase and some more uppercase letters with the same glyph/shape as latin ones.
- heavenlyblue 6y agoYeah, and those similar glyphs are not available on their own language keyboard.
- wongarsu 6y agoI frequently type German with a US layout with dead keys (so I can type "a to get ä). I also imagine that most Turkish developers type English on a Turkish layout, since Turkish contains all characters used by English.
- 6y ago
- mapgrep 6y agoDumb question, if you really need the exact string “info” in a given context, why not hard code it? What does .lower() or even a map liked the linked one actually buy you?
- nicoburns 6y agoPresumably it's for normalising input. Following the principle that you ought to be permissive in what data you accept, and strict in what data you give out.
- simion314 6y agoMaybe the input is case insensitive, for example if you work with html you might see "DIV","div" who knows some crazy dev or tool might generate "DIv" or "dIv" so is simpler to lowercase the input then work on it.
- ramses0 6y agoObTurkeyTest: http://www.moserware.com/2008/02/does-your-code-pass-turkey-test.html http://www.moserware.com/2008/02/does-your-code-pass-turkey-...
- Macha 6y ago07/04/2008 -> April 7th seems about as reasonable a result as July 4th, especially when you've explicitly opted in to a Turkish locale. I don't agree with the article's assertion that the format being interpreted according to the user's locale is wrong here, the one wrong part is a US centric programmer's expectation that PP-QQ-YYYY is an unambiguous format. Use YYYY-mm-dd when you need a format that's not ambiguous
- heavenlyblue 6y ago> PP-QQ-YYYY is an unambiguous format “US centric” is one way to say it
- snthd 6y ago> 07/04/2008 -> March 7th I think you mean April.
- Macha 6y agoFixed
- frabert 6y agoYYYY-mm-dd also plays nice with lexicographic ordering, which is why I always use it when I need to put dates in e.g. filenames
- Macha 6y agoI'm a European working primarily with Americans. My home country uses dd/mm/YYYY (or dd/mm for short) and the US uses mm/dd/YYYY for with mm/dd for short. I've switched to YYYY-mm-dd simply for my own sanity and if I omit the year I write the month in text format, such as "5 June".
- withinboredom 6y agoThe US military uses the almost same convention (dd-mmm-yyyy) so 07-aug-2020.
- TazeTSchnitzel 6y agoThe PHP interpreter has an internal reimplementation of string case conversion that's ASCII-only in order to avoid this problem.
- asddubs 6y agodoesn't php have this exact problem with their case-insensitive (hate that btw) function/method names and turkish localization? or did they actually fix it at some point?
- dhosek 6y agoI'm guessing that they might have "fixed" it by implementing the ascii-only tolower function, but yes, PHP used to not work properly with Turkish localization.
- TazeTSchnitzel 6y agoWhy do you think the interpreter needs such a function?
- chihuahua 6y agoI remember running into problems with SQL stored procedures where column and table names were case-insensitive, so you don't know if you've properly typed all the column and table names. Until a customer in Turkey eventually installs it and you find out you've missed the proper capitalization of an identifier containing the letter "I", and the stored procedure fails.
- heavenlyblue 6y agoThis is what I usually think about whenever people say yay to Unicode in language identifiers.
- formerly_proven 6y ago"I" is in ASCII.
- a1369209993 6y ago"İ" and "ı" are not.
- Pxtl 6y agoHonestly, I'm very pro case-insensitivity, but my experience with SQL servers have impressively demonstrated how not to do it. For example, MS SqlPackage, used for deploying schema, is case-insensitive... But that also means changes to text constants within your stored procs do not get treated as changes.
- mbostleman 6y agoHence toUpper/toLower is not a strategy that passes the Turkey Test for case insensitivity.
- formerly_proven 6y agoITT calling setlocale or std::locale::global(...) is ALMOST ALWAYS a heinously bad idea and should rarely be done, because it breaks tons of code (notably everything that uses printf/scanf and everything using stringstream).
- 60secz 6y agoStringly typed: Play stupid games, win stupid prizes.
- bayindirh 6y agoWelcome to the Turkish language, where we have ı, i, I and İ. In our language the conversion is as follows: - i <-> İ - ı <-> I We love our dots and preserve them. For a more detailed read, please see: https://blog.codinghorror.com/whats-wrong-with-turkey/ https://blog.codinghorror.com/whats-wrong-with-turkey/
- Natsu 6y agoAs I understand it, Turkish is one of the more important locales to test with because of things like this.
- bayindirh 6y agoTurkish is the only language which has the ı & I pair. Similarly, AFAIK, Turkish is again the only language with ğ and ş letters. So, by testing for Turkish, you test for a lot of European languages at once. Moreover we share some modified letters(ç, ü) with other Central European languages. If your program can pass “The Turkish Test”, you pass a lot of others too.
- anticensor 6y agoAzerbaijani too. Moreover, Azerbaijani has an additional letter ə, which sounds like /æ/.
- therein 6y agoI love the feeling of camaraderie arising from that partial mutual intelligibility of Turkish and Azerbaijani. That connection through language goes a long way. müqəddəs bacı millət :)
- deleted 6y ago[deleted]
- 1-more 6y agoPoor encoding can lead to the odd murder too: http://gizmodo.com/382026/a-cellphones-missing-dot-kills-two-people-puts-three-more-in-jail http://gizmodo.com/382026/a-cellphones-missing-dot-kills-two... > The use of "i" resulted in an SMS with a completely twisted meaning: instead of writing the word "sıkısınca" it looked like he wrote "sikisince." Ramazan wanted to write "You change the topic every time you run out of arguments" (sounds familiar enough) but what Emine read was, "You change the topic every time they are fucking you" (sounds familiar too.)
- iforgotpassword 6y agoWouldn't converting to nfkd/c first solve this issue too? My understanding of those forms was that they're made exactly for this case.
- jwilk 6y agoNo, these are ASCII strings, so they are already normalized.
- iforgotpassword 6y agoOh, I haven't used python much, but I thought it's all Unicode? If this were ascii it would work out of the box since there is no dotless lowercase i in ascii.
- estebank 6y agoThere are no code point for TURKISH LOWERCASE DOTTED I not for TURKISH UPPERCASE DOTLESS I, which means that the text doesn't carry enough information for roundtrip preservation. I believe this has proven to be a mistake but I'm not an expert. I don't know why it wasn't done.
- brewmarche 6y agoCase mapping and case folding are independent of normalization (in practice and it is the case here, see the end of SpecialCasing.txt) There is a good Unicode FAQ on the topic: < http://unicode.org/faq/casemap_charprop.html http://unicode.org/faq/casemap_charprop.html > E: to elaborate, I'm not sure whether the independence of case handling and normalization is guaranteed anywhere, and if we for example were to change the uppercase of ſ to something else than S then its compatibility forms' (s) case handling would differ. In practice the SpecialCasing.txt is designed to "make it work" (e.g. ſ uppercases to S).
- cazim 6y agohttp://www.moserware.com/2008/02/does-your-code-pass-turkey-test.html http://www.moserware.com/2008/02/does-your-code-pass-turkey-... This is old but still valid reading...
- garydgregory 6y agoSee also https://garygregory.wordpress.com/2015/11/03/java-lowercase-conversion-turkey/ https://garygregory.wordpress.com/2015/11/03/java-lowercase-...
- stevoski 6y agoFor a similar reason, Java on Mac and Linux was briefly broken for anyone using it in the Turkish locale. It was because in the Turkish locale, !“POSIX”.toLowerCase().equals(“posix”). Relevant bug report here: https://bugs.openjdk.java.net/browse/JDK-8047340 https://bugs.openjdk.java.net/browse/JDK-8047340
- jwilk 6y agoLooks like it's no longer the case in Python 3: Python 3.7.3 (default, Jul 25 2020, 13:03:44) [GCC 8.3.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> from locale import * >>> setlocale(LC_ALL, 'tr_TR.UTF-8') 'tr_TR.UTF-8' >>> 'INFO'.lower() 'info'
- xyst 6y agoPython 3.7.5 (default, Nov 5 2019, 22:30:48) [Clang 11.0.0 (clang-1100.0.33.12)] on darwin Type "help", "copyright", "credits" or "license" for more information. >>> from locale import * >>> setlocale(LC_ALL, 'tr_TR.UTF-8') 'tr_TR.UTF-8' >>> 'INFO'.lower() 'info' >>> '🧘 ️'.lower() '🧘\u200d️' >>> exit() There's something wrong with emojis + lower() though
- Dylan16807 6y agoIt lowercased the 'show this as emoji' variation selector to zero width joiner?
- anderskaseorg 6y agoOddly, it also wasn’t the case for Python 2 Unicode strings (u'INFO'), only for Python 2 byte strings ('INFO'). So it’s possible that Python 3 lost this behavior by accident.
- price 6y agoOn some more digging through history, it looks like the change in behavior for byte strings was intentional: https://github.com/python/cpython/commit/6ccd3f2dbcb98b33a71ffa6eae949deae797c09c https://github.com/python/cpython/commit/6ccd3f2dbcb98b33a71... Author: Guido van Rossum <guido@python.org> Date: Tue Oct 9 03:46:30 2007 +0000 Replace all (locale-dependent) uses of isupper(), tolower(), etc., by locally-defined macros that assume ASCII and only consider ASCII letters.
- jaclaz 6y agoOnly for the record, there is something very similar that may happen when creating CD/DVD's (please read when using mkisofs and similar), with the "dash" that when "capital" becomes underscore (but not only ) depending on the reference ISO 9660/Joliet/RockRidge convention in use. https://web.archive.org/web/20151007005513/http://www.911cd.net/forums//index.php?showtopic=25612 https://web.archive.org/web/20151007005513/http://www.911cd....
- kentonv 6y agoHas anyone here ever had a use case for toLower() where they actually wanted localization to apply? It seems to me that in practice, it's extremely rare to want to change case of real, natural-language text. When I have natural-language text, it's just a blob to me, and I don't want to touch it. The only time I ever want to lower-case or capitalize something, I'm working with identifiers meant for computer -- not human -- consumption. Usually, specifically, I'm dealing with identifiers that have annoyingly been defined to be case-insensitive even though the only humans that ever see them are programmers and programmers hate case-insensitivity. HTTP headers are a common example. I mostly write C++, and I end up writing code like: for (char& c: str) { if ('A' <= c && c <= 'Z') c = c - 'A' + 'a'; } Later on, some well-meaning developer on my team will come along and say "Ugh what is this NIH syndrome?" and then they "clean it up" as: #include <ctype.h> for (char& c: str) { c = tolower(c); } And then I have to say NOOOOOOO DON'T DO THAT YOU HAVE NO IDEA WHAT tolower() REALLY DOES! I struggle to imagine any real use case where you'd actually want locale-dependent tolower() other than, maybe, a word processor -- but if you're writing a word processor, you're probably not going to be depending on the language's built-in string APIs to do your text manipulation.
- rkangel 6y agoThis is a classic case of a 'why' code comment being needed. It's obvious what you're doing, but without a 2 line explanation, it's not clear why.
- kentonv 6y agoYeah I probably wrote that comment the first few times I did this but it's hard to write it the 50th time. Maybe I should have my own tolower() function that I can call so I only have to write the comment once but it just feels ridiculous somehow.
- Natsu 6y agoIt's far more ridiculous to repeat yourself over and over instead of making a simple function that describes exactly what you want and why.
- maweki 6y agoAs it isn't yet mentioned: for these cases the Python standard library explicitly has https://docs.python.org/3.8/library/stdtypes.html#str.casefold https://docs.python.org/3.8/library/stdtypes.html#str.casefo... (str.casefold), which aggressively lowercase-normalizes strings with an algorithm from the unicode standard. Every case comparison using lower() instead of casefold() can be considered a bug.
- Alex3917 6y ago> Every case comparison using lower() instead of casefold() can be considered a bug. If you just casefold two strings and compare them, it's still a bug. You need to normalize them to NFKC first.
- pas 6y agoIs NFKC necessary, isn't NFKD enough? (As in you have to normalize and decompose both strings, but at that point you can check them for equality, and doing the canonical composition isn't needed, right?)
- Alex3917 6y agoI think that would work if you're just checking for equality and want to minimize processing. I guess as a web developer I always just assume people are going to be storing strings in a database after normalizing them, so would want to minimize string length.
- FrontAid 6y agoChanges to the casing might also change the value's length. E.g. uppercasing the German ß will transform it to SS. Example using JavaScript: 'ß'.toUpperCase(); // returns 'SS' https://en.wikipedia.org/wiki/%C3%9F https://en.wikipedia.org/wiki/%C3%9F
- schoen 6y agoThere is apparently a multi-decade controversy about that: https://en.wikipedia.org/wiki/Capital_%E1%BA%9E https://en.wikipedia.org/wiki/Capital_%E1%BA%9E (with German language authorities recently endorsing the idea that ß can have a distinctive uppercase form "ẞ")
- deleted 6y ago[deleted]
- dathinab 6y agoWhich can be both correct and wrong depending on context. Normally there is no such thing as a capital ß, so it was decided that if for some unreasonable reason you do uppercase it you go with SS. But then for some all-caps usages this is not right. E.g. a all caps name of an restaurant as placed above the restaurants door. In which case it was common to have a ß in a all-caps name like FOOßBAR. So they decided that for reasons like this we now have an (EDIT: semi?) official uppercase ß. So all in all this and other examples in other languages mean you should never do a case insensitive comparison by upper/lower casing both sides, it won't work reliable.
- seqizz 6y agoYeah, there were some weird bugs about that. I remember one in a media player. Also "info".upper() would be İNFO probably.
- tryauuum 6y agoUnrelated story about Russian language. The first letter of russian alphabet is А, the last one is Я. So it's natural to try to match russian words with '[А-Яа-я]+'. But this is a recipe for disaster, this regexp doesn't match words with 'Ё' in them like "Артём". This is due to the fact that regexp ranges work on byte values. All letters of russian language have neatly ordered byte values, except for the Ё.
- blkhawk 6y agonot that unusual - for German for instance üöäÜÖÄß need to be added so all words can be matched.
- a3w 6y agonow, there is even a capital ß ;)
- Forge36 6y agoOut of curiosity I tried on my phone: ß Ss SS So my phone doesn't have that yet!
- maxpro 6y agomine has it - ẞ the small one is ß
- Igelau 6y agoIs the capital supposed to be shorter?
- zaarn 6y agoIt's wider. Depends on your font and it's support though.
- a1369209993 6y ago
- decafbad 6y agoPlease stop doing this. Don't bind lower() upper() functions to environment variables or anything else system related. Sun did this in Java and doesn't even bother to mention the issue in documents. It caused huge problems for more than a decade. You can just make string lowercase() uppercase() function work the same everywhere, regardless of locale settings. Provide a special case function lowercaseTR() or so. This works very well in Go. By the way, Azerbaijan has the same problem because they accepted help from wrong guys when they switched to Latin.
- netsharc 6y ago> lowercaseTR() Huh, that works well if we know the input string is in Turkish. What if this information is not available as you're writing the code? And what will lowercase()/uppercase() be hard coded to do, and what are they supposed to output when the input isn't ASCII?
- decafbad 6y agoGive me an example. I'll try to find the best -IMHO- solution.
- price 6y agoYou'll be glad to hear that Python did stop doing this: Python 3 has never behaved this way, and its `lower` and `upper` methods have always been independent of your locale or anything else from your system. The workaround in the OP was added in 2006 (note the reference to an issue on "SF", i.e. SourceForge -- another era!), and is now long obsolete.
- decafbad 6y agoVery much so. Thanks.
- sedatk 6y agoNote to the next language designer: don't use strings as a substitute for enums.
- crazygringo 6y agoSerious question. Why on earth would you hard-code these, instead of simply call a lowercase function in the en-US locale? These are English words. Naively lowercasing them according to whatever locale the server or user has set seems like a terrible programming practice. Any call to a lowercase function should be explicitly including an argument that specifies it's English, no? In the same way we've all learned to never store times without an explicit timezone (even if it's UTC), or locate a string offset without knowing your encoding... you should never perform language transformations (case changes, accent removal, etc.) without a locale. Hardcoding these things is just patching over the symptoms without addressing the cause, no?
- beeforpork 6y agoMy genius idea was once to use toupper() to normalise paths on Windows, which are case-insensitive. One day, a customer from Azerbaijan reported that my application failed to access a file in C:\WİNDOWS\...
- tryauuum 6y agoi feel your pain
- CodesInChaos 6y agoIn the Danish locale "aa" doesn't start with "a".
- alkonaut 6y agoRepeat after me: don’t do string operations without explicit locale. Don’t do string operations without explicit locale. I don’t know why so many languages have string functions that should take a locale but provide an overload that doesn’t and which uses the system locale as the default. It can’t be what many developers actually want, yet it has become the norm. Worse, code using a default locale appears to work on the developers machine and in production, until someone parses a number in France or lowercases a string in Turkey, which is a late and expensive discovery of the bug. The default shouldn’t be the system locale, it should be an invariant locale. And I’ll go so far as arguing this invariant locale should be invariant across systems (meaning it can’t just defer to a system C library either).
- madeofpalk 6y agoI ran into this with C#/.NET on Windows - I tried to convert a string "1.3" to the float 1.3, and it failed on languages that use comma as their decimal separator. That was a learning experience.
- alkonaut 6y agoIndeed. As a person from a comma country, I find these mistakes in most code bases I look at. It makes it frustrating to contribute to open source, for example. Perhaps it’ll make you feel better about your parsing bug that even the C# compiler (Roslyn) code base had several of these issues.
- layer8 6y ago> I don’t know why so many languages have string functions that should take a locale but provide an overload that doesn’t and which uses the system locale as the default. That‘s a relict from the past, before Unicode became prevalent, where systems used to ever only work in a single locale, and where users expected applications (running locally of course) to use the local system locale. Hence applying the system locale to everything was the standard behavior for applications. The C standard library was defined in that way, and since then every other runtime (usually based on C at some level) does the same.
- 6y ago
- dependenttypes 6y agoAh yes, locales. Everyone loves them https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f027338b0fab0f5078971fbe https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
- shagmin 6y agoI learned about this in javascript when I discovered Angular has its own lowercase method. Apparently it's internal only now. https://github.com/angular/angular.js/commit/1daa4f2231a89ee88345689f001805ffffa9e7de https://github.com/angular/angular.js/commit/1daa4f2231a89ee...
- maple3142 6y agoI think things like these should be explicit. Even it is convenient to have a default, it should be what most people would expect. For example, instead of .lower(), we can have .lower_ascii(), .lower_turkish() or .lower(locale) . But I know it would be tedious to use if you need to specify it everytime, so it makes sense to have a .lower(locale=DEFAULT_LOWER_LOCALE) . As for what should DEFAULT_LOWER_LOCALE be, it is worth debating, but I think it shouldn't introduce unexpected behavior.
- baybal2 6y agoThe "İ" strikes again
- mro_name 6y agowhat a gorgeous source-comment. Makes the non-obvious crystal-clear.
- dusted 6y agoI think we should have stopped at ASCII, I don't care that my language has letters not in there, it'd be neater if we just did now like back then: "This is a computer, so everything is in English" :) Or adapt the alphabet to use ASCII.
- lxe 6y agoThe practice of converting enum-like keys into their string representation by using toString, toLower, etc seems convenient but gets very contrived very fast. How do you deal with underscores? What about using the message in a sentence? I say, use the enum in your code as a conditional or something but always explicitly write out the messages intended for the user.
- paledot 6y agoI'm going to be a bit controversial here and say that that mapping logic should always exist even if toLower() were reliable across all locales. You're mapping between different use cases here, eg. internal to logfile to API to database to method name to whatever, and inserting magic transformations in your constant values rather than treating them as different tokens for different use cases constrains you and introduces unnecessary amounts of "magic".