3 ms·
The Unicode Technical Standard [1] (different from Unicode Standard) recommends treating identifiers (filenames, variable or fuction names, email adresses, user
by cdirkx 7y ago
The Unicode Technical Standard [1] (different from Unicode Standard) recommends treating identifiers (filenames, variable or fuction names, email adresses, usernames, etc.) different from normal text. There is a special class of 'identifier characters' which already exludes a lot of sneaky characters like invisible punctuation and obscure scripts that are not in modern use.
Additionally there are 5 additional restriction levels for identifiers depending on your specific situation:
1. ASCII only
2. single script
3. single script or Latin+{Japn, Hanb, Kore}
4. single script or Latin+{any, excluding Cyrillic, Greek}
5. Not containing any characters in the recommended blacklist of characters for use in secure contexts
6. No restrictions other than normal identifier restrictions
For any specific strings that might be confused, there is an algorithm to compute the visual 'skeleton' of a string and match it against that of another string to test if they are confusable.
[1] https://www.unicode.org/reports/tr39/ https://www.unicode.org/reports/tr39/