3 ms·
UTF-8 is a nifty compression scheme - like a Huffman encoding that doesn't require a table to implement, but the committee doesn't deserve any credit for that (
by zerofan 10y ago
UTF-8 is a nifty compression scheme - like a Huffman encoding that doesn't require a table to implement, but the committee doesn't deserve any credit for that (unless Thompson and Pike were on the committee, which I doubt).
Besides, the reason most of us actually like UTF-8 is because it leaves ascii alone (which is all I ever use) while pretending to handle the general case. It doesn't help end users or programmers deal with any of the nonsense around multiple ways to encode glyphs (combining codes vs accented codes), deal with surrogates (yes, people encode surrogates in UTF-8), lexical sorting, or anything else. I'll bet there are dozens of incompatible ways strings are UTF-8 encoded in the real world, each of them a bug for interoperability, and all of that blame falls on Unicode being a terrible standard.
So yes, I'm sure.