4 ms·
Well if memory cost wasn't an issue then UCS-4 would be nicest. But 4x the memory for most strings is currently unacceptable.
by TwoBit 5y ago
Well if memory cost wasn't an issue then UCS-4 would be nicest. But 4x the memory for most strings is currently unacceptable.
- ohazi 5y agoThere are a lot of instances where I wouldn't mind the memory cost, but would very much mind not automatically being compatible with ascii strings.
- lokedhs 5y agoThat is true, but the benefits of UTF-32 is minor compared to UTF-8, and might not be worth the cost. And I say that as someone who is developing a language which only has UTF-32 support. The problem is that even with UTF-32, doing things like splitting strings is inherently unsafe, so you are still going to need a Unicode library to do proper splitting by grapheme cluster. In practice, almost all string splitting works on ASCII text, and assumes everything else is data that should not be manipulated. For this, UTF-8 is perfectly acceptable.
- fanf2 5y agoUCS-4 is still a variable-length encoding, because various accented characters and emoji use multiple code points. One advantage of UTF-8 is that it makes you confront variable-length characters head on.