3 ms·
You need raw values for serialization over the wire and to files. And for any kind of compression. And for searching (it's hard to build indexes in Unicode beca
by nullwasamistake 7y ago
You need raw values for serialization over the wire and to files. And for any kind of compression. And for searching (it's hard to build indexes in Unicode because of uneven symbol widths).
You could convert to a different character format but at some point you need raw values. If you want to be able to read text created by something else (text editor, browser, different OS) you have to expose raw access.
- lmm 7y ago> You need raw values for serialization over the wire and to files. And for any kind of compression. You need to encode as bytes with a particular encoding. That's not the same thing as "getting the raw value". > And for searching (it's hard to build indexes in Unicode because of uneven symbol widths). It may be hard but it's important to do it right. If searching for "café" doesn't find "caf◌́e" then you've got a problem. > You could convert to a different character format but at some point you need raw values. If you want to be able to read text created by something else (text editor, browser, different OS) you have to expose raw access. Again, no. You need to be able to decode byte sequences that have particular encodings into strings of characters. That doesn't mean you need to expose the internal representation of those strings/characters.
- nullwasamistake 7y ago> Again, no. You need to be able to decode byte sequences that have particular encodings into strings of characters. You keep reiterating this, but it's not feasible. To avoid exposing the raw values, the language would need to support all possible encodings.
- lmm 7y agoNonsense. What is it that you can do with a "raw value" that you can't do with the representation in a particular encoding? I mean, if you wanted to import a character that doesn't have a unicode codepoint then you couldn't decode that character from UTF-8 - but if the language is built with no support for non-unicode characters then even if you did have access to the internal representation of a character, that wouldn't help you (e.g. the language's built-in character functions for things like checking the case of the character won't handle a non-unicode character properly).
- nullwasamistake 7y agoIn Java at least, you can write your own Charset implementation then the language will support it normally. This uses the byte raw access I'm ranting about to work
- lmm 7y agoYes and no - a Java Charset is something that can convert from a buffer of bytes to a buffer of utf16 code units. In Java that happens to be the internal representation of a String, but it doesn't have to be - you just need built in support for encoding/decoding a string as utf16. A simple proof of this is that you can write custom character sets in Python too, even though there's no way to have raw byte access to a Python character (because it's different on different platforms/builds).
- kuschku 7y agoYou could ask the API to get the UTF-8 or UTF-16 encoding. You could ask the API to get the unicode character id. But how it’s stored internally, you’d never care
- nullwasamistake 7y agoI had to convert some strings to Base85 some years ago to integrate with an ancient system. There's no way to support all the different character sets out there "out of the box". Unless you want a crippled language you need to support access to raw values so people can write their own handler
- kuschku 7y agoWhat did you convert to Base85? The UTF-8 representation? The Unicode character ids, as 32-bit, appended after another (UTF-32)? The UTF-16 encoding? Big or Little Endian? Or with BOM? There are no raw values. The same string, used in Qt/C++ vs. stdlib/C++ results in different raw values, because it’s encoded and handled differently. The fact that you believe there is a single true "raw value" of a string already shows why it’s important to make this explicit, because your code will break if used in different environments, and definitely won’t return the correct result.