4 ms·
My understanding is that UTF-8 is not a good representation for non-european alphabets. So do you think UTF-8 is always the best internal string representation
by diroussel 5y ago
My understanding is that UTF-8 is not a good representation for non-european alphabets.
So do you think UTF-8 is always the best internal string representation? Or just for English speakers?
For Mandarine what would be optimal?
- rectang 5y agoMandarin is an interesting case. Most of the Han characters used by Mandarin fall within the basic multilingual plane and thus occupy 2 bytes in UTF-16 but 3 bytes in UTF-8. However, for web documents, most markup is ASCII which is only one byte. So for Mandarin web documents, the space requirements for UTF-8 and UTF-16 are about a wash. When you add in interoperability concerns, since so much text these days is UTF-8, for Mandarin at least UTF-8 is a perfectly defensible choice. (A harder problem is Japanese — Japan really got screwed over with Han unification, so choosing Shift-JIS over any Unicode encoding is often best.) FWIW I covered the space requirements of various encodings and various languages in this talk for Papers We Love Seattle: https://www.youtube.com/watch?v=mhvaeHoIE24&t=39m14s https://www.youtube.com/watch?v=mhvaeHoIE24&t=39m14s
- amelius 5y agoI genuinely wonder: is the space requirement of text encodings really an important issue in this age of large photo and video content?
- astrange 5y agoMemory for strings is often more important because there's a lot more of them, and image memory can be file-backed more often but strings need to be swapped to disk.
- jcranmer 5y agoLooking at my browser's memory reporting, strings take up about ~2-3% of total memory usage, most of which is probably ASCII. If it were UCS-2, that would make it ~5% of total memory usage, and UCS-4 ~10%. That's small numbers, but as a whole-program impact, it's significant enough to motivate performance engineers to actually try to compress those strings down a bit.
- amelius 5y agoIt depends on how strings are counted. If every String object is atomized/interned then e.g. the string "div" is stored once, but on a 64bit system you have 8 bytes for a pointer and another 8 bytes for bookkeeping things such as length.
- FridgeSeal 5y agoIf you ever want to do something stuff with text efficiently (full text search and associated processing) I’d argue that it’s quite important.
- Dylan16807 5y agoIf you want to do search, you need collation, and you can't use any standard encoding for that data.
- rectang 5y agoIt's important to distinguish between sorting in code point order and sorting according what a user would expect for their language. However, sorting in code point order (which is actually equivalent to sorting by memory comparison for UTF-8) is enough to build an inverted index data structure, commonly used for fulltext search. And just like FridgeSeal asserted, the memory footprint of the text representation has performance implications for such an application. (Source: I wrote a search engine library.)
- tsimionescu 5y agoWill that handle things like matching accented and unaccented characters? If I search for 'stefan' in a text that makes frequent references to 'Ștefan', will it correctly find those matches?
- catblast01 5y ago> (A harder problem is Japanese — Japan really got screwed over with Han unification, so choosing Shift-JIS over any Unicode encoding is often best.) This statement needs more support. I think “screwed over” is a bit harsh, since I’m not aware the impact on Japanese was anymore than the rest of CJK. Despite the Han unification controversy, Unicode has been heavily adopted in Japan. The space requirements are basically the same as all CJK. Half-width kana is heavier since they are one byte in shift jis but they’re relatively uncommon.
- ChrisSD 5y agoAs far as I'm aware one problem is displaying text. Japanese readers generally need a Japanese font to correctly display Japanese text if it's Unicode. This becomes a problem when you potentially have text that can come from different languages. E.g. a Japanese font will display Chinese incorrectly. On the web you can work around this using the lang attribute to tell the browser how text should be interpreted. It's notable that, for example, traditional and simplified Chinese does not have this problem because they are encoded separately. Another problem is missing characters. Some people have complained of not being able to write their own name. I'm not sure to what extent this has been solved through Unicode updates.
- majewsky 5y agoI have been learning Japanese for about a year now, so I don't have that much experience reading Japanese text yet [1]. I'm aware of some of the visual differences between Chinese and Japanese fonts, but I have not yet had trouble reading Japanese text set in a Chinese font. If you have any specific examples for Kanji that are difficult to recognize in a Chinese font, I'd be interested. [1] Although on the other hand you could argue that I'm spending more conscious effort reading Japanese text than a native speaker would.
- ChrisSD 5y agoSorry, I can't speak confidently on this because I can't read Japanese (or Chinese or Korean). I can only report what I've been told by users. Including language metadata with text was felt to be especially important for Japanese and Korean users. I was told the difference was like having "595 kg" displayed as "5P5 kg". That is, it's possible to decipher the intended meaning but it looks wrong and it takes a moment to work out what was meant. Depending on the language some glyphs can be mirror images, have extra strokes, strokes missing or in different places or at different angles.
- klodolph 5y agoSo, the advantage of UTF-16 is that CJK text will use 33% less space. Does this mean that “UTF-8 is not a good representation for non-European alphabets?” It may be less efficient but the difference does not seem shocking to me, considering that for most applications, the storage required for text is not a major concern—and when it is, you can use compression.