3 ms·
The Perl doc mostly refers back to suggestions from UAX#31, along with specific notes from UTR#36 and UTS#39. The suggestion it cares most about from UAX#31 is
by ubernostrum 8y ago
The Perl doc mostly refers back to suggestions from UAX#31, along with specific notes from UTR#36 and UTS#39.
The suggestion it cares most about from UAX#31 is to restrict to scripts that actually are in use (i.e., don't let people name variables using Linear B characters). If you want to layer that on top of a base pool of identifier characters taken from XID_Start/XID_Continue, you can, and that's not the same as "invent your own base pool of identifier characters".
The big thing you get from reading UTR#36 and UTS#39 is learning how to detect or prevent homoglyph attacks (like people registering "paypal.com" but with Cyrillic "a").
And UTS#39 takes the ideas from UAX#31 all the way and gives you an example of defining profiles on top of the base set of identifier characters to deal with specific issues. It's OK to do that!
Finally, the things that, in a programming language's allowed identifier syntax, would be prevented by going to a more restrictive profile on top of XID_Start/XID_Continue, are vanishingly rare and tend to be exploitable only if you're already completely owned. For example, if someone can slip a mixed-script confusable identifier into your source, they can already slip things into your source; you're owned in so many ways at that point that it starts seeming silly to focus with laser intensity on just this one issue. Which means that if you just want an easy-ish to implement baseline recommendation, XID_Start/XID_Continue isn't that bad.