5 ms·
utf8 vs utf8mb4 bit me in one of the most frustrating bugs I've ever had to debug. At Scribophile members can write critiques for people's writing, with commen
by acabal 5y ago
utf8 vs utf8mb4 bit me in one of the most frustrating bugs I've ever had to debug.
At Scribophile members can write critiques for people's writing, with comments inserted inline. The underlying software dates back to the PHP5 days when MySQL only had the utf8 option. Everything was working fine for years, when all of a sudden users started complaining that from time to time, they would submit a critique and it would be mysteriously cut off at random places.
The problem was very intermittent, didn't happen very often, and there was no seeming rhyme or reason to it. But when it did happen, it was catastrophic because members would lose hours of work, seemingly at random!
All kinds of testing scaffolding and logging was put in place to try to find the problem with no luck. Then, after quite some time, we realized what the problem was.
At that time, emoji keyboards were brand new; most phones/tablets had limited support, and people didn't yet use phones and tablets for everything like they do now. But they were out there. Some users who were using these new emoji keyboards were inserting emoji smiley faces as they were writing critiques. In Scribophile's web interface, everything looked fine; but when the user submitted the critique, MySQL tried to insert a multibyte Unicode character into a regular utf8 field, and SILENTLY threw that character and all of the data after it away!!
Boy were we upset at MySQL about that one. But at least we figured it out!
- BoxOfRain 5y agoIt's stories like this that stop me complaining online about a long day of futile debugging!
- gpderetta 5y agoYou probably mean outside of the basic multilingual plane. Any non-ascii character is multibyte in utf8.
- webtopf 5y agoI had the pleasure of finding that same behaviour in an application around 2015. When someone commented on a food order "Extra cheese please :folded hands: but no shrimp, I'm allergic." But one thing I haven't found out is why some emojis did go through. The basic ones like a simple :smile: it seemed to me. Could it be that some only need 3 bytes and when more and more emojis got released, they went into 4 bytes? Edit: HN stripped the emojis from my comment... I would put a rolling eyes emoji here if I could.
- bombcar 5y agoExactly that - three byte emojis work fine in utf8 but 4 byte ones need utf8mb4 - and the four byte ones are the new ones that support the "color variations" and other similar things. Apple adding those probably caused more upgrades to utf8mb4 than any amount of pleading from languages that actually needed them ever would have.
- kingcharles 5y agoI looked through the master emoji list and I can't see that any expand into the 4th byte, but this stuff is REALLY complicated and I have no idea. https://unicode.org/Public/emoji/14.0/emoji-test.txt https://unicode.org/Public/emoji/14.0/emoji-test.txt
- chrismorgan 5y agoI believe you’re confusing two distinct concepts: Sounds like you’re looking at how many scalar values are in an extended grapheme cluster (e.g. U+1F635 U+200D U+1F4AB dizzy face + zero width joiner + dizzy symbol, that’s one “character” made up of three scalar values). The thing in question here is UTF-8 code units (bytes), and how many UTF-8 needs to encode a single scalar value. UTF-8 needs four bytes for anything above U+FFFF, and almost all emoji are above that.
- kingcharles 5y agoI'm confused because U+1F635 can be stored in 3 bytes? I see nothing above 0xFFFFFF which would need the 4th byte? Or am I being dense?
- kangalioo 5y agoIn UTF-8, only some bits are used for the actual character codes. The rest are control bits for the decoder to know where a character starts and ends. https://en.m.wikipedia.org/wiki/UTF-8#Encoding https://en.m.wikipedia.org/wiki/UTF-8#Encoding
- riffraff 5y agoI hit the same issue in some app I wrote. It's what convinced me never to use MySQL ever again.