6 ms·
Yea, CJK (Chinese, Japanese, Korean) breaking is particularly complex. Google has done a lot of work, and have this open source implementation which uses NLP. I
by LikeAnElephant 6y ago
Yea, CJK (Chinese, Japanese, Korean) breaking is particularly complex. Google has done a lot of work, and have this open source implementation which uses NLP. It's the best I've personally come across:
https://github.com/google/budou https://github.com/google/budou
- oehtXRwMkIs 6y agoCan't imagine it being difficult for Korean.
- Asooka 6y agoKorean is especially difficult. Chinese uses only hanzi and you have a limited set of logogram combinations that result in a word. Japanese is easier because they often add kana to the ends of words (for e.g. verb conjugation) and you can use that (in addition to the Chinese algorithm) to delineate words - there is no word* where a kanji character follows a kana character. Korean on the other hand uses only phonetic characters with no semantic component, so you just have to guess the way a human guesses. * with a few exceptions, of course :)
- qiqitori 6y agoKorean uses spaces.
- oehtXRwMkIs 6y agoCurious where you came up with "phonetic characters with no semantic component". It's just an alphabet with spaces. With each block representing a syllable. It's easier than Latin.
- nine_k 6y agoYes, Hangul is one of the two scripts known to me designed with.proper logic and regularity, which makes it easy to use. (The other is Tengwar.)
- NikolaeVarius 6y agoOld korean lacked spaces. Modern korean uses spaces
- tasogare 6y agoYour comment is a mix of a few right things and lot of wrongs, so I’ll give more information so the casual reader doesn’t form an incorrect opinion based on it. - Korean segmentation is way easier than Chinese and Japanese because it uses spaces between words (they are thus distinctive, not like Vietnamese which use it on syllable boundaries. Vietnamese consequently also requires segmentation) - Chinese and Japanese segmentation are hard NLP problems that are not fixed, so they are in no way "easier" than the same take for other languages - The limited valid combination of characters that form words in Chinese doesn’t mean segmentation is easy because there is still ambiguity in how sentence can be split. There is still no tool that produce "perfect" result - difference in scripts is indeed used in some segmentation algorithms for Japanese, but that doesn’t solve the issue totally - the phonetic/non-phonetic parts of Chinese characters, have been used in at least one researcher paper (too lazy to find the reference again, it didn’t worked well anyway) but are not in state of art method. So contemporary Korean not using a lot of Hanja anymore has no influence on the difficulty of segmenting it
- polm23 6y agoI don't speak Korean, but my understanding is that some applications don't just use spaces because of how they're used in loan words. The Korean for "travel bag" has a space in it but you might want it as one token, for example. There's a fork of MeCab that has some different cost calculations related to whitespace for use with Korean.
- oehtXRwMkIs 6y agoThat's interesting, but it would be nice to see the actual example. Not sure why you would use the loan word for travel bag, or what Hangul you're talking about exactly.
- polm23 6y agoSorry, "travel bag" is just an example I remember someone mentioning before. You can see the Mecab fork with example output here, but the docs are all in Korean so I can't really follow it. https://bitbucket.org/eunjeon/mecab-ko-dic/src/master/ https://bitbucket.org/eunjeon/mecab-ko-dic/src/master/
- jcampbell1 6y agoI wrote a Chinese Segmenter that is available on the web: https://chinese.yabla.com/chinese-english-pinyin-dictionary.php?define=%E6%88%91%E6%98%AF%E7%94%B5%E8%84%91%E7%A8%8B%E5%BA%8F%E5%91%98 https://chinese.yabla.com/chinese-english-pinyin-dictionary.... It does basic path finding, and then picks the best path based on the following rules: 1) Fewest words 2) Least variance in word length (e.g. prefer a 2,2 character split vs a 3-1 split) 3) Solo Freedom (this is based on corpus analysis which tags characters with a probability of being a 1 character word. For example 王家庭 (this is either "Wang Household" (王 家庭) or "Prince's courtyard" (王家 庭) and we split as Wang Household, because Wang 王 is a common name that frequently appears in isolation, and 庭 is less likely to be in isolation. It is interesting that solo freedom works better than comparing the corpus frequency of "Prince" 王家 vs "Household" 家庭. It works reasonably well. A surprising number of people use it every day.
- thaumasiotes 6y ago> the corpus frequency of "Prince" 王家 What? 王家 doesn't mean "prince". Or at least, there is no such dictionary entry in the ABC dictionary or in the 汉语大词典. I would expect 王家 to mean "prince's household", in the same way that 皇家 means "imperial household". 汉语大词典 has two glosses for 王家: 1. 犹王室,王朝,朝廷。 [Equivalent to "royal family"/"royal court".] 2. 王侯之家。 [An aristocratic household.] There's no problem with the concept of the phrase 王家庭 meaning "prince's courtyard", since a home can easily contain a courtyard. But the phrase should arguably be segmented 王-家-庭. (Or not -- there's very little to distinguish the idea 'one word, "the prince's household"' from 'two words, "prince"/"household"'.) Regardless of that choice, the courtyard is being associated with a household, not a person.
- numpad0 6y agoSo to illustrate the situation the input would be like “Royal, House, Hold” and whether it’s supposed to be “Royal-Household” or “Mr. Royal’ household” or “Royal House, hold ... ” is up to context right
- thaumasiotes 6y ago
- krackers 6y agoI don't think their own lexing backend is actually open-source, as Budou just relies on a choice of 3 backends (MeCab, TinySegmenter, Google NLP) to do the lexing. I'm assuming Google NLP performs the best, but that isn't free and certainly not open source.
- Wowfunhappy 6y agoDoes this mean that highlighting doesn't work properly when you're offline?
- krackers 6y agoI think the implementation used in Chrome is different from the ones used in Budou. The implementation in Chrome is dictionary based as one of the parent threads mentioned, and that is completely open-source albeit probably doesn't produce as good a result as their homegrown NLP stuff.