4 ms·
> Afaik, most Asian characters can be encoded as 2 bytes in UTF-8 I overstated my case for sure. The Unified ideograms are four words wide, but the non-unified
by guns 16y ago
> Afaik, most Asian characters can be encoded as 2 bytes in UTF-8
I overstated my case for sure. The Unified ideograms are four words wide, but the non-unified extensions are larger, iirc. And while disk space is hardly a problem anymore, you might see why programmers from 10 years ago may have made different choices.
> But it still seems to me that extending Unicode to also include those variatons of characters is a better and more clean solution than these workarounds.
I certainly can't argue with that. But Unicode is a standard, and real problems don't have time to wait around for standards bodies. In the eyes of many East Asian organizations, Unicode is broken now, and so the burden falls on the programmer.
Even here in the US, there are tons of data sitting around in tables encoded in Windows-1251 and ISO-8859-1. Having had to deal with UTF-8 and Latin-1 mismatches in the past, I don't find ruby1.9's encoding all that onerous myself.