3 ms·
As I said, there are many reasons UTF-8 is a better encoding. And indeed compact, backwards compatible, encoding of ASCII is one of them.
by ChrisSD 5y ago
As I said, there are many reasons UTF-8 is a better encoding. And indeed compact, backwards compatible, encoding of ASCII is one of them.
- glandium 5y agoIt is less compact than UTF-16 for CJK languages, FWIW.
- maxdamantus 5y agoThat's only really true when the entire string is using CJK characters, but as far as I know, that's going to be fairly rare. For example, this Chinese Wikipedia page is almost twice as big in UTF-16 as in UTF-8, because the vast majority of the content seems to be HTML tags: $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | wc -c 1877366 $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | iconv -f utf-8 -t utf-16 | wc -c 3345724 If I strip out all the printable ASCII characters, it does indeed become about 33% more compact in UTF-16: $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | wc -c 311226 $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | iconv -f utf-8 -t utf-16 | wc -c 213444 With gzip compression it is only about 10% more compact: $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | gzip -9c | wc -c 98373 $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | iconv -f utf-8 -t utf-16 | gzip -9c | wc -c 89031 So I guess if you're archiving pure CJK text, maybe you could get a 10% benefit, though I suspect non-Unicode encodings of that text would be more compact anyway.
- bigbizisverywyz 5y agoMicrosoft did a comparison between UTF-8 and UTF-16 for Sql Server and found pretty much the same results. I can't find the more detailed article, but this summarizes it: https://techcommunity.microsoft.com/t5/sql-server-blog/introducing-utf-8-support-for-sql-server/ba-p/734928 https://techcommunity.microsoft.com/t5/sql-server-blog/intro... From what I can remember, UTF-8 consumes more CPU as it's more complex to process, has space savings for mostly ascii & European codepages, but can significantly bloat storage sizes for character sets that consistently require 3 or 4 bytes per character.