3 ms·
The C standard defines the execution character set. First it defines the base execution character set. That includes: upper and lower case A to Z, digits 0 to
by jsmith45 3y ago
The C standard defines the execution character set.
First it defines the base execution character set. That includes: upper and lower case A to Z, digits 0 to 9, these symbols !"#%&'*+,-./:;<=>?[\]^_{|}~ , space, horizontal tab, vertical tab, form feed, alert, backspace, CR, LF. These will all have single byte encoding.
The full execution character set is further defined as the basic execution characters (single byte), plus additional locale-specific characters (either single or multi-byte):
> The source character set may contain multibyte characters, used to represent members of the extended character set. The execution character set may also contain multibyte characters, which need not have the same encoding as for the source character set. For both character sets, the following shall hold:
> - The basic character set shall be present and each character shall be encoded as a single byte.
> - The presence, meaning, and representation of any additional members is locale-specific. [ED: emphasis added]
> - [allows for shift-dependent encodings]
> - A byte with all bits zero shall be interpreted as a null character independent of shift state. Such a byte shall not occur as part of any other multibyte character.
---
Which is to say, that the basic characters must have identical encodings in all supported char* strings, and everything else is dependent on the locale, which means setting the locale effectively modifies the execution character set. If you want to correctly interpret any string literal that exceeds the basic character set, you need to know what locale the compiler used when converting to the execution character set, and set the locale to match.
Any encoding that does not have the base characters having the same byte values as are not technically valid as a string, but could be represented as an arbitrary byte array, even if you choose to spell that byte array as `char*`. This also implies that no locale may legally use such as encoding.