3 ms·
I think that moving the world away from HTML and XML as markup languages is likely to be considerably harder than simply defaulting to using UTF-8 in any new co
by lambda 11y ago
I think that moving the world away from HTML and XML as markup languages is likely to be considerably harder than simply defaulting to using UTF-8 in any new code for which there isn't already a natural encoding to choose based on the platform or APIs you're developing on.
Furthermore, even beyond the markup, the vast majority of written content is in the Roman alphabet, of which most characters are in the ASCII range, and those that aren't fit into 2 bytes of UTF-8. 55% of text on the web is in English, then languages using the Roman script or other scripts in the 2-byte UTF-8 range make up 21 of the next 25 most commonly used languages on the internet; of the top 25 languages, only Chinese, Japanese, Korean, and Thai are in ranges that require 3 bytes in UTF-8 for the majority of their characters (https://en.wikipedia.org/wiki/Languages_used_on_the_Internet https://en.wikipedia.org/wiki/Languages_used_on_the_Internet).
So for the vast majority of text that is processed, UTF-8 is dramatically more efficient than a hypothetical UTF-24 or the real UTF-32. Now, you might say "well, it should all compress away anyhow", but when dealing with text in RAM, it is generally not compressed (and that would make handling it far more complex than just using UTF-8), and memory capacity, latency, and bandwidth can all be important.
Trying to deal with text as fixed-width characters is simply incorrect. Almost every non-trivial means of handling text needs to support arbitrary length strings, needs to deal with clusters of more than one codepoint as a single unit, and needs to iterate over the characters linearly at least once (and can then store byte offsets for random access later on). Dramatically increasing storage requirements of text for the non-benefit of being able to deal with fixed-width codepoints just doesn't make sense.