9 ms·
I had the pleasure of finding that same behaviour in an application around 2015. When someone commented on a food order "Extra cheese please :folded hands: but
by webtopf 5y ago
I had the pleasure of finding that same behaviour in an application around 2015. When someone commented on a food order "Extra cheese please :folded hands: but no shrimp, I'm allergic."
But one thing I haven't found out is why some emojis did go through. The basic ones like a simple :smile: it seemed to me. Could it be that some only need 3 bytes and when more and more emojis got released, they went into 4 bytes?
Edit: HN stripped the emojis from my comment... I would put a rolling eyes emoji here if I could.
- bombcar 5y agoExactly that - three byte emojis work fine in utf8 but 4 byte ones need utf8mb4 - and the four byte ones are the new ones that support the "color variations" and other similar things. Apple adding those probably caused more upgrades to utf8mb4 than any amount of pleading from languages that actually needed them ever would have.
- kingcharles 5y agoI looked through the master emoji list and I can't see that any expand into the 4th byte, but this stuff is REALLY complicated and I have no idea. https://unicode.org/Public/emoji/14.0/emoji-test.txt https://unicode.org/Public/emoji/14.0/emoji-test.txt
- chrismorgan 5y agoI believe you’re confusing two distinct concepts: Sounds like you’re looking at how many scalar values are in an extended grapheme cluster (e.g. U+1F635 U+200D U+1F4AB dizzy face + zero width joiner + dizzy symbol, that’s one “character” made up of three scalar values). The thing in question here is UTF-8 code units (bytes), and how many UTF-8 needs to encode a single scalar value. UTF-8 needs four bytes for anything above U+FFFF, and almost all emoji are above that.
- kingcharles 5y agoI'm confused because U+1F635 can be stored in 3 bytes? I see nothing above 0xFFFFFF which would need the 4th byte? Or am I being dense?
- kangalioo 5y agoIn UTF-8, only some bits are used for the actual character codes. The rest are control bits for the decoder to know where a character starts and ends. https://en.m.wikipedia.org/wiki/UTF-8#Encoding https://en.m.wikipedia.org/wiki/UTF-8#Encoding
- deleted 5y ago[deleted]
- chrismorgan 5y agoU+1F635 is encoded in UTF-8 as bytes {f0, 9f, 98, b5}.
- kingcharles 5y agoThank you, that explains it!
- Dylan16807 5y agoTo give a slightly more complete answer as to why, utf-8 was designed to be both backwards compatible with ASCII and self-synchronizing. So you can't use half the values of a byte when encoding code points outside of ASCII, and the remaining 128 values need to be split between starting bytes and continuation bytes. At best you're going to fit about 6 bits into each byte, and unicode spans the equivalent of just over 20 bits, so you need 4 bytes max. utf-8 opts to keep encoding very simple, and in 2-4 bytes it can store 11, 16, and 21 bits respectively. If you discarded ASCII compatibility, you could create a 1-3 byte self-synchronizing unicode encoding. But at that point maybe you should be compressing your text instead. Or you could weaken the compatibility and create "utf-1 but better" to fit into 3 bytes.
- geoduck14 5y ago>Apple adding those probably caused more upgrades to utf8mb4 than any amount of pleading from languages that actually needed them ever would have. I think there is a lesson here - maybe I'll find it after my coffee