3 ms·
The model of a Java String, as presented via the API, is as an array of 16-bit unsigned integers. Strings use a UTF-16-like encoding to represent non-BMP Unicod
by smarks 3y ago
The model of a Java String, as presented via the API, is as an array of 16-bit unsigned integers. Strings use a UTF-16-like encoding to represent non-BMP Unicode characters. The main difference from UTF-16 is that illegal sequences of 16-bit values (such as unpaired surrogates) are permitted.
The "UTF-8 by Default" refers to the default character set used for decoding and encoding, when Java Strings are read and written. Prior to JEP 400, the default character set was "platform specific." It was often UTF-8, but sometimes it was not, in which case hijinks ensued.
As noted elsewhere the internal representation of Java Strings is either UTF-16-like or is Latin-1, if all the characters of the String can be encoded in Latin-1. We think about using UTF-8 as an internal representation constantly. Doing so would potentially save space and reduce decoding/encoding overhead.
Unfortunately doing this is difficult. The key issue is that tons of code out there treats a String as a UTF-16 array (since that's what's in the API) so there are coding patterns like this:
for (int i = start; i < end; i++) {
char ch = str.charAt(i);
// do something with ch
}
If the internal representation were UTF-8, simplistically, `charAt(i)` would change from O(1) to O(N). Of course there are cleverer things one could do, such as converting to UTF-16 lazily. Now `charAt()` allocates memory and has variable latency. Well then partial conversion could be done and the current iteration point could be cached. This might work, but now String has state, and it's thread-specific as well. Etc.
- Phrodo_00 3y ago> If the internal representation were UTF-8, simplistically, `charAt(i)` would change from O(1) to O(N). 16 bits is not enough space to store all unicode characters. Does that mean that java's charAt won't join chars for codepoints over U+FFFF? Edit: Yeah, that's exactly what happens[1]. That's not a very nice implementation at all . > Because 16-bit encoding supports 2^16 (65,536) characters, which is insufficient to define all characters in use throughout the world, the Unicode standard was extended to 0x10FFFF, which supports over one million characters. The definition of a character in the Java programming language could not be changed from 16 bits to 32 bits without causing millions of Java applications to no longer run properly. To correct the definition, a scheme was developed to handle characters that could not be encoded in 16 bits. > The characters with values that are outside of the 16-bit range, and within the range from 0x10000 to 0x10FFFF, are called supplementary characters and are defined as a pair of char values. [1] https://docs.oracle.com/javase/tutorial/i18n/text/unicode.html https://docs.oracle.com/javase/tutorial/i18n/text/unicode.ht...
- Spivak 3y agoPython also has this problem which you only find out once you're bitten by it. The encoding of all the built-in file manipulation machinery is platform specific so you either have to do `sys.setdefaultencoding("utf-8")` or add the encoding to every open call that uses text mode. So kudos to every language that axes platform dependent encodings. Honestly platform dependent anything is so annoying to deal with. I would love if languages make it painfully explicit when you depend on platform defaults.
- rtpg 3y agoEven after Py2 -> Py3, the platform-dependent decoding sitting around for _so long_ was so annoying. Just constantly running into issues with file reads because servers were configured a certain way. "Just specify the encoding" is like... fine, but annoying when it's some third party script I just can't use. I'm at least glad we had some guarantees from being Python 3.
- rtpg 3y agoDoes Java have any sort of precedent for dynamically swapping out implementations at the first call of charAt to a more specialized string class? I have to imagine there's a lot of string stuff that _doesn't_ do this at all.
- daniel_grady 3y ago> Now `charAt()` allocates memory and has variable latency. It’s Java, right? Everything is variable latency.