4 ms·
Unicode in identifiers is an exceptionally bad idea Unicode identifiers can be perfectly well-defined, and many languages have them. The general idea is an ide
by ubernostrum 8y ago
Unicode in identifiers is an exceptionally bad idea
Unicode identifiers can be perfectly well-defined, and many languages have them. The general idea is an identifier can start with any character that has derived property XID_Start, and the remainder can be any characters that have derived property XID_Continue.
different characters looking identical, or the same character encoded in different ways
Follow the recommendation of UAX #31 and normalize identifiers to NFKC prior to comparing them. For example, the ligature variants in the article should not -- if recommendations are followed -- be different identifiers. Here's some Python (3):
>>> import unicodedata
>>> raw = ['vpnTrafficPort', 'vpnTrafficPort', 'vpnTrafficPort', 'vpnTrafficPort']
>>> normalized = set(unicodedata.normalize('NFKC', s) for s in raw)
>>> normalized
{'vpnTrafficPort'}
So there wouldn't be a "surprise" lurking in Python -- all four strings are legal identifiers, but all four of them are also the same identifier (because Python applies normalization).
The real difficulty with Unicode identifiers is in places where you really can't avoid Unicode: user inputs. Those who don't read Unicode technical reports are doomed to suffer the moment they build, say, a user-account system that comes into contact with the real world.
When you allow Unicode identifiers, you don't "internationalize" your code
No, you let other people localize their code. We (the tech world) spent decades making everyone else learn English as a prerequisite to learning a programming language. Now we have the ability to lessen that burden and let the non-ASCII world (which does in fact include English!) spend less time learning English and more time writing code. We probably should do that, even if it seems icky to you.
- continuational 8y agoSo instead of having to learn English as a prerequisite to understand code bases, you propose that we learn a dozen or more different natural languages instead? Can you imagine what it would be like if every library had to be released in N different translations? Sounds like an exceptionally bad idea. A huge step back for our field.
- ubernostrum 8y agoIf you take the position that every codebase everywhere must be available to and understood by every programmer everywhere, sure, you have to settle on a common language. Luckily, I don't think anybody takes that position, or at least takes it seriously. But there are plenty of monolingual dev teams out there whose shared language isn't English and isn't written in ASCII. Why should they be forced to write code in English if nobody who works with it is an English speaker? I already regularly see people post questions to Stack Overflow and mailing lists and IRC where names of classes and functions and variables are ASCII but not English (i.e., someone writes a class called "Utilisateur", not "User"). Why not let them just use their language fully?
- hyperpape 8y agoThis is a false dichotomy. English will continue to be the lingua franca of programming for the foreseeable future, and full professional competence will probably require being able to read APIs that are in English. However, for the sake of learners everywhere, it is important that it be possible to program using non-ascii identifiers. Asking a 10 year old who doesn't speak English to learn what "while" means is quite different from requiring their variables to be in a language they don't understand.
- kazinator 8y agoBut ASCII is not English. If identifier are restricted to ASCII, then what that requires is romanization, not specifically English. char *speicher; // German for memory (define f (keizoku) ...) ; Japanese for continuation
- adwn 8y ago> Follow the recommendation of UAX #31 and normalize identifiers to NFKC prior to comparing them. You're proposing that we replace the rules for identifier equality (comparing case-sensitive ASCII strings composed of A-Z, a-z, 0-9, _ is very easy to understand) with a complex system which hardly anyone can remember in its entirety, and even fewer people will want to learn? All so that according to you, some people need to learn a little less English? I sure hope that those normalization rules have been translated to lots of languages, otherwise those people would still need to learn English... > Now we have the ability to lessen that burden and let the non-ASCII world (which does in fact include English!) spend less time learning English and more time writing code. Do you seriously think that a programmer needs lower English skills just because they can use non-ASCII characters in their function names? That programming language's standard library is in English, its documentation is in English, external libraries are in English, their documentation is in English, StackOverflow is in English, but you need to "spend less time learning English" because you can name your variable "über_schwellwert" instead of "exceeds_threshold"?! Come on.
- squaresmile 8y agoI think it would be great if people don't have to learn English to program. There are tons of people out there using computers, programming with very very poor English knowledge. Besides technical implementation issues, I see no reason to oppose allowing them to name things in a language they are more familiar with. Of course you can keep using ASCII strings for your program but they should be able to name things however they like too.
- adwn 8y ago> I think it would be great if people don't have to learn English to program. But that's the point – you'll still need to learn just as much English when you have Unicode identifiers as when you only have ASCII identifiers.
- hyperpape 8y agoThe rules for unicode normalization only make a difference when you use them. As an English language programmer, they'll probably never affect you, and you're free to use your linter to require all code in your projects use only ASCII identifiers, just to prevent any potential problems. Perl apparently handles this well: you have to declare that a file uses a particular script if you're using non-ASCII identifiers.
- kazinator 8y ago> Follow the recommendation of UAX #31 and normalize identifiers to NFKC prior to comparing them That's insane. Not in my programming language (lawn, driveway, ..).
- hyperpape 8y agoSimply following the unicode recommendations will leave you with some weird issues: https://github.com/rust-lang/rust/issues/4928 https://github.com/rust-lang/rust/issues/4928 https://twitter.com/aisamanra/status/923346798093090816 https://twitter.com/aisamanra/status/923346798093090816
- ubernostrum 8y agoThe rejection of XID_Start/XID_Continue linked from the Rust issue is not, as far as I can tell, due to "weird" issues -- it's because the sets of characters with those properties can change over time as Unicode adds more characters. Other languages tend to solve this by stating that version X of the language uses identifiers having those properties in version Y of the Unicode database. This is a problem for C and C++ because of how infrequently those languages update their standards and how glacially slowly implementations adopt new versions of the standards. It's less of a problem for languages like Rust where the language spec and standard toolchain evolve on a much faster (at least one update per year) cadence. The tweet you link complains about U+2800. That character is not in XID_Start or XID_Continue. http://www.unicode.org/Public/11.0.0/ucd/DerivedCoreProperties.txt http://www.unicode.org/Public/11.0.0/ucd/DerivedCoreProperti... So a language adopting XID_Start/XID_Continue would not allow U+2800 in an identifier. And in fact if I try it in Python I get "SyntaxError: invalid character in identifier".
- hyperpape 8y agoThanks, I hadn’t realized that covered U+2800—my mistake. I haven’t fully digested the links to the C working group, but as those Perl docs point out, you have to do some work even after restricting to XID_Start/Continue, even if you don’t have to forbid other characters.
- ubernostrum 8y agoThe Perl doc mostly refers back to suggestions from UAX#31, along with specific notes from UTR#36 and UTS#39. The suggestion it cares most about from UAX#31 is to restrict to scripts that actually are in use (i.e., don't let people name variables using Linear B characters). If you want to layer that on top of a base pool of identifier characters taken from XID_Start/XID_Continue, you can, and that's not the same as "invent your own base pool of identifier characters". The big thing you get from reading UTR#36 and UTS#39 is learning how to detect or prevent homoglyph attacks (like people registering "paypal.com" but with Cyrillic "a"). And UTS#39 takes the ideas from UAX#31 all the way and gives you an example of defining profiles on top of the base set of identifier characters to deal with specific issues. It's OK to do that! Finally, the things that, in a programming language's allowed identifier syntax, would be prevented by going to a more restrictive profile on top of XID_Start/XID_Continue, are vanishingly rare and tend to be exploitable only if you're already completely owned. For example, if someone can slip a mixed-script confusable identifier into your source, they can already slip things into your source; you're owned in so many ways at that point that it starts seeming silly to focus with laser intensity on just this one issue. Which means that if you just want an easy-ish to implement baseline recommendation, XID_Start/XID_Continue isn't that bad.
- kazinator 8y ago> The real difficulty with Unicode identifiers is in places where you really can't avoid Unicode: user inputs. Those who don't read Unicode technical reports are doomed to suffer the moment they build, say, a user-account system that comes into contact with the real world. Simply reject any piece of input, such as a user display name, which uses characters from two or more different scripts, or purely typographical devices such as ligatures.
- ubernostrum 8y agoNot quite. Reject any mixed-script confusable (and for certain types of inputs, like email addresses, validate sub-components separately -- someone might get to choose their local-part but not their domain). For strings which aren't mixed-script confusable, if you need to use them as identifiers and perform comparisons with them, normalize to NFKC and case fold. Doing so will eliminate any ligatures, stylistic variants, composed versus decomposed forms, or other things that people use to call Unicode "weird".
- deleted 8y ago[deleted]