6 ms·
For portable code, use UTF-8 only, and only the part in ASCII. This also helps to distinguish digits from alphas etc. Unicode is a different code. (and as I l
by hakre 3y ago
For portable code, use UTF-8 only, and only the part in ASCII. This also helps to distinguish digits from alphas etc.
Unicode is a different code.
(and as I learned today, Julia code might be exclusive here)
- Macha 3y agoUTF-8's using only characters from the ASCII subset is just ASCII, this was kind of the point of UTF-8 at inception.
- hakre 3y agoYes, exactly, and Unicode has more encodings than UTF-8 so stating it explicitly makes clear which encoding of Unicode that is. Writing only ASCII would not be relaying on Unicode, but on ASCII only. IMHO a bit shortsighted today.
- Macha 3y agoUsing UTF-8 but only the ASCII bits, and using ASCII are the exact same activity, however. I don't see how the first is any different to the latter, other than which document you look up to tell you that 00110000 is "0". The difference between an "ASCII subset of UTF-8" parser and ASCII parser is whether the error message on encountering a high bit set to 1 is "Invalid character" or "Character not in permitted ranges". If your point is that your program should use Unicode internally, this is already true for most programming languages but is independent of your input, plenty of languages routinely converting UTF-16 to UTF-8 to work with Linux or the web when they use UTF-16 internally, or UTF-8 to UTF-16 to work with Windows when they use UTF-8 internally. But I'd argue if you've done that work already, what harm is there in allowing your users to type € or á or 風?
- hakre 3y ago> Using UTF-8 but only the ASCII bits, and using ASCII are the exact same activity, however. Which was my point. The encoding might be different. For ASCII 7 bit is fine, for UTF-8 and only ASCII, 8 bits are required and, as you also already point out, the high bit must be set to zero with it. So the point is to name the encoding (UTF-8) of the ASCII characters. And the input was portable code, not user input of a program.