3 ms·
> Again, no. You need to be able to decode byte sequences that have particular encodings into strings of characters. You keep reiterating this, but it's not fe
by nullwasamistake 7y ago
> Again, no. You need to be able to decode byte sequences that have particular encodings into strings of characters.
You keep reiterating this, but it's not feasible. To avoid exposing the raw values, the language would need to support all possible encodings.
- lmm 7y agoNonsense. What is it that you can do with a "raw value" that you can't do with the representation in a particular encoding? I mean, if you wanted to import a character that doesn't have a unicode codepoint then you couldn't decode that character from UTF-8 - but if the language is built with no support for non-unicode characters then even if you did have access to the internal representation of a character, that wouldn't help you (e.g. the language's built-in character functions for things like checking the case of the character won't handle a non-unicode character properly).
- nullwasamistake 7y agoIn Java at least, you can write your own Charset implementation then the language will support it normally. This uses the byte raw access I'm ranting about to work
- lmm 7y agoYes and no - a Java Charset is something that can convert from a buffer of bytes to a buffer of utf16 code units. In Java that happens to be the internal representation of a String, but it doesn't have to be - you just need built in support for encoding/decoding a string as utf16. A simple proof of this is that you can write custom character sets in Python too, even though there's no way to have raw byte access to a Python character (because it's different on different platforms/builds).