5 ms·
Internally, the model of a String is UTF-16-like, a sequence of unsigned 16-bit `char` values. However, it's not strictly UTF-16 because a Java String permits u
by smarks 4y ago
Internally, the model of a String is UTF-16-like, a sequence of unsigned 16-bit `char` values. However, it's not strictly UTF-16 because a Java String permits unpaired surrogates and invalid code points.
Externally, Java has historically dealt with whatever character set was in use on whatever platform it was running on. As time as gone on, most platforms have converged on UTF-8 for external data storage and interchange. Time for Java to do that too, at least by default.
- jillesvangurp 4y agoThe common mistake is relying on any 'default' for this. There's no such thing. This has caused me a huge amount of headaches over the years in the form of messed up strings in data exports, databases, or API outputs. There's always one place where somebody forgets to specify Utf-8 and relies on the default. Which is often not Utf-8 on lots of operating systems depending on how they are configured. Even Linux has issues with this. Even in 2023. Most distributions set a default encoding of UTF-8 of course but if you start your server early in the boot process as root, none of that may be applied. If you write your own systemd scripts, be sure to make this explicit as well. I've found this out the hard way where the default was latin-1 and combined with somebody not doing what they should be doing, you end up dealing with corrupted data. It's why not specifying a encoding when decoding bytes to a string is a mistake in Java that things like spotbugs check for and warn against. It's one of those things that IMHO should have been deprecated in Java ages ago. Instead they keep on adding more APIs that do the same thing. Removing it would probably break lots of code. But the point is that that code should already be considered broken and should be fixed. Thankfully, Kotlin uses sensible default parameters in Kotlin extension functions that do similar things for this (of Utf-8). If you want something else, you have to be explicit about it. That's the only sane way to deal with character encodings IMHO. As long as you don't use Java APIs directly, things are fine.
- PaulHoule 4y agoI think the industry is moving towards UTF-8 everywhere as a default however the platform is misconfigured. Python certainly is. I think it is rare but not unheard of to find files in a different encoding and if that is the case it is not deliberate but an accident because of a misconfigured default encoding.
- jillesvangurp 4y agoThe industry moved more than two decades ago already. However operating systems continue to do the wrong thing for no good reason other than that they have been doing that since forever. UTF-8 was already a thing when Java launched in 1995. It was proposed as a standard three decades ago. As I mentioned, I've been dealing with weird UTF-8 related issues throughout my career. As soon if you have APIs that rely on the default and operating systems that don't set that to UTF-8, you end up with corrupted content. There are all sorts of ways to make this happen: APIs that set a mime type but no character encoding, progams that open and save files, databases that don't default to using utf-8 (or like mysql actually have some non standard 3 byte version of utf-8 that doesn't deal with emojis and a lot of other characters). Linux distributions don't set UTF-8 as the character encoding until late in the boot process. IMHO there's no good reason why it should ever do anything else than UTF-8 in a shell from the microsecond it boots until it shuts down. Changing all that is hard however. Not for technical reasons but for reasons of backwards compatibility with systems that are arguably part of the problem here. I think the root cause for this might be a blind spot that people in English speaking countries have since they don't use a lot of special characters and therefore don't see the problem. As soon as you start using Spanish, Swedish, German, etc. You are in trouble. And forget about any non latin scripts. Latin-1 is not appropriate for any of those. Utf-8 is a better default that should be usable everywhere. Some countries might prefer utf-16. But that's mainly a memory optimization that should not matter much in 2023.
- Someone 4y ago> Internally, the model of a String is UTF-16-like, a sequence of unsigned 16-bit `char` values. That is not 100% true anymore. Java 9 (September 2017) changed “the internal representation of the String class from a UTF-16 char array to a byte array plus an encoding-flag field. The new String class will store characters encoded either as ISO-8859-1/Latin-1 (one byte per character), or as UTF-16 (two bytes per character), based upon the contents of the string. The encoding flag will indicate which encoding is used.” (https://openjdk.org/jeps/254 https://openjdk.org/jeps/254)
- smarks 4y agoThe model of a String, internal to the JVM, as manifested in the `java.lang.String` API, is still UTF-16-like. The internal representation has changed to be either UTF-16-like or Latin-1. JEP 254 didn't change the API. This is invisible to Java programs, and they can't tell which internal representation is being used. If a Java program wants a byte array from a String, it still has to go through the APIs that involve codeset conversion. There are suitable fast paths, though. For example, if the program requests Latin-1 and the internal representation is Latin-1, the result is just a straight copy of the internal array.