4 ms·
> So one Unicode character can be up to 5 bytes long and take up the same canvas space as 3 characters. 5 bytes? In what encoding?
by rcoveson 3y ago
> So one Unicode character can be up to 5 bytes long and take up the same canvas space as 3 characters.
5 bytes? In what encoding?
- dredmorbius 3y agoThere was an emoji example of length seven posted to HN recently: <https://news.ycombinator.com/item?id=36159443 https://news.ycombinator.com/item?id=36159443>
- jychang 3y agoEmojis with skin color, mostly
- rcoveson 3y agoNo, that's a ZWJ sequence. Those can be arbitrarily long. Doesn't explain where "5 bytes" comes from.
- throwaway2037 3y agoIs there a maximum number of "zero-width joiner" (ZWJ) sequences that can be combined to create a single-width emoji? (Yes, I know the term "single-width" is a loaded term.) I cannot find a precise answer.
- lelanthran 3y ago> 5 bytes? In what encoding? I believe UTF-8 reserved up to six bytes for a single character.
- rcoveson 3y agoYes, UTF-8 was up to 6 bytes per character early on. Some broken implementations like MySQL's limit it to up to 3 bytes per character. The actual number is 4. So what is "up to 5"?
- chipsa 3y agoI think with decomposed Hangul, you can end up with 6 or more bytes per character, due to each part of it being two bytes, and 2-4(?) parts per character.
- WorldMaker 3y agoSome "single character" emoji easily exceed 5 bytes in all encodings. You may think ZWJ sequences are cheating, but emoji isn't the only language encoded in Unicode with complex ZWJ sequences.
- rcoveson 3y agoMy question is about the phrase up to 5. What in Unicode is up to 5? Codepoints are up to 4 in all the encodings I know. ZWJ sequences may as well be arbitrarily long. What is "up to 5"?
- cstrahan 3y agoThe original quote for reference: > So one Unicode character can be up to 5 bytes long and take up the same canvas space as 3 characters. FWIW, I didn't read that as suggesting an upper bound of 5 bytes, but rather as an example using arbitrary numbers: N bytes of code units could, depending on the font providing the glyph(s) for the respective grapheme(s), could be rendered at M times the size of, say, the letter A, where N != M -- despite the font otherwise being monospaced. Which is just another way of saying that you must consult the font for the character widths involved. I think you're reading that quote as an assertion that: For any grapheme G, G can be encoded in at most 5 bytes. While what I think was being said was: There exists a grapheme G, where G is encoded in 5 bytes, and the respective glyph happens to be displayed at 3 times a single character (e.g. the letter A), despite the font otherwise being monospaced. Therefore you *must* consult the font for each glyph to correctly determine character widths.
- rcoveson 3y agoContinue to the next sentence of context: > You also need to read ahead as there are combination characters, for example a smiley combined with the color brow becomes a brown smiley. Emphasis mine. Clearly combination characters are being treated separately. Frankly I think it's crazy to read "up to 5 bytes" and not think that it suggest an upper bound. I think you're reaching for a highly questionably interpretation of a totally unambiguous clause. If the author meant to express what you're saying, they would certainly have written: "Some Unicode characters are 5 bytes long and take up the same canvas space as 3 characters". Which would still look incorrect if they followed it with the sentence "You also need to read ahead as there are combination characters...". It is far more likely that the author is simply mistaken and should have said 4 bytes, and perhaps used the word "codepoint" instead of "character" in the original sentence. That's a perfectly understandable technical error, while the reinterpretation you're putting together would imply an error of colloquial language.