6 ms·
Adding a variable decorator/annotation like @Unicode(german,french) would be a good stop-gap. You could only use ASCII characters unless you specified the scrip
by Dagonfly 5y ago
Adding a variable decorator/annotation like @Unicode(german,french) would be a good stop-gap. You could only use ASCII characters unless you specified the script that you want to use. One could even set a max limit on how many scripts per variable. Because while I have used German characters in variables before (only if I'm referring to some law or spec), I never had a use case for more than 2 scripts within one variable.
- lifthrasiir 5y agoFor your information the relevant Unicode specification is the Script_Extensions property [1]. (You can't easily filter by languages, so you should filter by scripts.) [1] https://www.unicode.org/reports/tr24/tr24-32.html#Script_Extensions https://www.unicode.org/reports/tr24/tr24-32.html#Script_Ext...
- est31 5y agoThe multiple scripts per variable thing is implemented in Rust via a lint. For the explicit enabling of single scripts, I have suggested that for Rust, but sadly people preferred allowing all identifiers (while giving an option to only have ascii but I'd argue this is unfair for anyone who only wants to use a specific non-ascii language, why do they have to suddenly allow all languages in their code base?). There are also practical concerns, like who says what a language is, which characters it contains, how that language is called, etc? Someone has to maintain all these lists.
- chrismorgan 5y ago> who says what a language is, which characters it contains, how that language is called, etc? The Unicode Consortium already maintains all of that data in the CLDR (Common Locale Data Registry).
- silvestrov 5y agoI think this is a good idea because once in a while you need to write non-ascii characters in names. This mostly comes up when implementing tax rules or government administrative divisions as some countries have names/concepts which have no good translation into English, so you are left with using the non-English name, which often contains non-ASCII characters.