4 ms·
"On the other end of the spectrum, languages like Chinese, Japanese, and Korean have thousands of characters." This is not exactly true. Chinese has an iconic
by denimboy 17y ago
"On the other end of the spectrum, languages like Chinese, Japanese, and Korean have thousands of characters."
This is not exactly true. Chinese has an iconic lexicon where each glyph is a single word. Both Cantonese and Mandarin speakers both use the same lexicon, but have different pronunciations. There are several lexicons (pinyin, big5, ancient) and thousands of glyphs in each.
Japanese has three lexicons; hiragana, katagana, and kanji. Kanji is the oldest and adapted from Chinese. Glyphs are iconic. Hiragana and katagana were developed in Japan and are phonetic. Together they form most of what you see as Japanese text today. I think katagana is used more for foreign, non-Chinese words. There is also romanji which is essentially English letters. Anyway, apart from kanji which is Chinese, there are less than 100 hiragana and katagana glyphs.
Korean is even simpler. They too borrowed from the Chinese and occasionally still use some Chinese glyphs, but the official lexicon is Hangul. Hangul is phonetic and has 24(?) basic glyphs. Some glyphs can be combined into compound glyphs called double consonants and double vowels making about 40 glyphs total. A Korean word can be written by breaking it down into syllables, combining glyphs to form a syllable super-glyph, then putting those together to form a word.
The explanation of unicode and python3 is great. I just wanted to clear up the misconception stated in the first paragraph. Nothing more to see here...
- chairface 17y agoI am having trouble seeing what you're claiming. Are you saying that there are not thousands of characters in Japanese and Korean because they borrowed from the Chinese? If so, I'd say your correction is misplaced. These characters must still be taken into account for a charset to be used for these languages. In any case, I found many of your comments to be irrelevant to the question of how many characters must be used in a language. For instance, the difference between Cantonese and Mandarin pronunciation doesn't have anything to do with this issue. Nor does Chinese origin. edit: I just spent a little time researching Korean (which I know much less of than Chinese or Japanese), and now I understand more what you were saying about it. However, it seems to me that each "super-glyph" as you call them counts as a character, as far as any charset is concerned. The fact that they can be broken up into constituent glyphs is irrelevant.
- bobbyi 17y agoReally? I thought the idea of Han Unification was that the duplicated characters between the CJK languages all map to the same unicode codepoints.
- chairface 17y agoYes that's true, but even so, you can't fit all those characters into 8 bits, which is basically the point of that first section. Also, I don't see how mapping to the same codepoint would mean that Chinese has thousands of characters, while Japanese does not. They just share many of those thousands in common.
- chairface 17y agoAlso, now that I am reading focusing on your words and not trying to figure out what you mean, a lexicon is a collection of words, not a collection of characters. The concepts overlap somewhat in the case of Chinese, but katakana, for instance, is certainly not a lexicon. (apologies for replying again, but the time for editing has passed)