3 ms·
> Does someone has any sources about this? Try this for example: http://tclab.kaist.ac.kr/~otfried/Mule/unihan.html http://tclab.kaist.ac.kr/~otfried/Mule/unih
by guns 16y ago
> Does someone has any sources about this?
Try this for example: http://tclab.kaist.ac.kr/~otfried/Mule/unihan.html http://tclab.kaist.ac.kr/~otfried/Mule/unihan.html
> Esp, why is SHIFT-JIS important? Is Unicode not capable of encoding the full Japanese language? What information would you loose by encoding with Unicode?
The problem is that SHIFT-JIS, and other regional east asian encodings, codify variations of characters that Unicode combine into one. The Japanese in particular dislike this conflation of characters that they feel should be separate. This process of combining characters was called Han Unification:
http://en.wikipedia.org/wiki/Han_unification http://en.wikipedia.org/wiki/Han_unification
Now imagine that you're a Japanese user who feels that three similar ideograms are in fact sematically different, but a standards body has decided that the differences are merely stylistic and has combined them into one form: I'm sure you wouldn't be as willing to lay your language at the feet of Unicode and swear allegiance to the one true encoding.
Also consider that UTF-8 is optimized for efficient encoding of Western languages: a japanese text may be in practice four to eight times the size of a sematically similar version in a western language. SHIFT-JIS presumably does not suffer from this problem.
> This seems to be a much better solution than this encoding nightmare. Esp in a language like Ruby, I want to have it simple and straightforward.
Except that languages are not really that simple or straightforward. I appreciate the flexibility of Ruby 1.9's encoding system. The only thing really broken right now are the tools that wycats mentions in this post.
- albertzeyer 16y agoThank you for the information and the links. But it still seems to me that extending Unicode to also include those variatons of characters is a better and more clean solution than these workarounds. In your link, the author speaks also about some fonts being incomplete. The right solution about this would be to fix the fonts, not to switch to another encoding with its own font. Having one encoding (Unicode) that is able to encode just everything would simplify everything. Btw., you are speaking about space efficiency. Afaik, most Asian characters can be encoded as 2 bytes in UTF-8. You cannot get it much better. And the space required for text is in most cases much smaller compared to other media. Also, if there is much redundancy in it, it can easily be compressed (also in some transparent way if needed, like all Ruby strings with more than 32kb are automatically compressed internally or so).
- guns 16y ago> Afaik, most Asian characters can be encoded as 2 bytes in UTF-8 I overstated my case for sure. The Unified ideograms are four words wide, but the non-unified extensions are larger, iirc. And while disk space is hardly a problem anymore, you might see why programmers from 10 years ago may have made different choices. > But it still seems to me that extending Unicode to also include those variatons of characters is a better and more clean solution than these workarounds. I certainly can't argue with that. But Unicode is a standard, and real problems don't have time to wait around for standards bodies. In the eyes of many East Asian organizations, Unicode is broken now, and so the burden falls on the programmer. Even here in the US, there are tons of data sitting around in tables encoded in Windows-1251 and ISO-8859-1. Having had to deal with UTF-8 and Latin-1 mismatches in the past, I don't find ruby1.9's encoding all that onerous myself.