5 ms·
The rejection of XID_Start/XID_Continue linked from the Rust issue is not, as far as I can tell, due to "weird" issues -- it's because the sets of characters wi
by ubernostrum 8y ago
The rejection of XID_Start/XID_Continue linked from the Rust issue is not, as far as I can tell, due to "weird" issues -- it's because the sets of characters with those properties can change over time as Unicode adds more characters.
Other languages tend to solve this by stating that version X of the language uses identifiers having those properties in version Y of the Unicode database. This is a problem for C and C++ because of how infrequently those languages update their standards and how glacially slowly implementations adopt new versions of the standards. It's less of a problem for languages like Rust where the language spec and standard toolchain evolve on a much faster (at least one update per year) cadence.
The tweet you link complains about U+2800. That character is not in XID_Start or XID_Continue.
http://www.unicode.org/Public/11.0.0/ucd/DerivedCoreProperties.txt http://www.unicode.org/Public/11.0.0/ucd/DerivedCoreProperti...
So a language adopting XID_Start/XID_Continue would not allow U+2800 in an identifier. And in fact if I try it in Python I get "SyntaxError: invalid character in identifier".
- hyperpape 8y agoThanks, I hadn’t realized that covered U+2800—my mistake. I haven’t fully digested the links to the C working group, but as those Perl docs point out, you have to do some work even after restricting to XID_Start/Continue, even if you don’t have to forbid other characters.
- ubernostrum 8y agoThe Perl doc mostly refers back to suggestions from UAX#31, along with specific notes from UTR#36 and UTS#39. The suggestion it cares most about from UAX#31 is to restrict to scripts that actually are in use (i.e., don't let people name variables using Linear B characters). If you want to layer that on top of a base pool of identifier characters taken from XID_Start/XID_Continue, you can, and that's not the same as "invent your own base pool of identifier characters". The big thing you get from reading UTR#36 and UTS#39 is learning how to detect or prevent homoglyph attacks (like people registering "paypal.com" but with Cyrillic "a"). And UTS#39 takes the ideas from UAX#31 all the way and gives you an example of defining profiles on top of the base set of identifier characters to deal with specific issues. It's OK to do that! Finally, the things that, in a programming language's allowed identifier syntax, would be prevented by going to a more restrictive profile on top of XID_Start/XID_Continue, are vanishingly rare and tend to be exploitable only if you're already completely owned. For example, if someone can slip a mixed-script confusable identifier into your source, they can already slip things into your source; you're owned in so many ways at that point that it starts seeming silly to focus with laser intensity on just this one issue. Which means that if you just want an easy-ish to implement baseline recommendation, XID_Start/XID_Continue isn't that bad.