8 ms·
This is often determine by Unicode and not the browsers specifically (though some browsers could override the suggested Unicode approach). Each unicode charact
by LikeAnElephant 6y ago
This is often determine by Unicode and not the browsers specifically (though some browsers could override the suggested Unicode approach).
Each unicode character has certain properties, one of which is whether that character indicates a break before / after itself.
I've done extensive research on this for my job, but unfortunately don't have time to do the whole writeup here. Here are several resources for those who are interested
Info on break opportunities:
https://unicode.org/reports/tr14/#BreakOpportunities https://unicode.org/reports/tr14/#BreakOpportunities
The entire Unicode Character Database (~80MB XML file last I checked)
https://unicode.org/reports/tr44/ https://unicode.org/reports/tr44/
The properties within the UCD are hard to parse, here's a reference if you're interested:
https://unicode.org/reports/tr14/#Table1 https://unicode.org/reports/tr14/#Table1
https://www.unicode.org/Public/5.2.0/ucd/PropertyAliases.txt https://www.unicode.org/Public/5.2.0/ucd/PropertyAliases.txt
https://www.unicode.org/Public/5.2.0/ucd/PropertyValueAliases.txt https://www.unicode.org/Public/5.2.0/ucd/PropertyValueAliase...
Overall, word / line breaking in Unicode in no-space languages is a very difficult problem. Where the UCD says there can be a line break isn't where a native speaker would put one. In order to do it correctly you have to bring in Natural Language Processing, but that has its own set of complexities.
In summary: I18N is hard!
- swang 6y agoYes. This seems to work even when you pass it Chinese while maintaining ja-JP as the language function tokenizeJA(text) { var it = Intl.v8BreakIterator(['ja-JP'], {type:'word'}) it.adoptText(text) var words = [] var cur = 0, prev = 0 while (cur < text.length) { prev = cur cur = it.next() words.push(text.substring(prev, cur)) } return words } console.log(tokenizeJA("今天要去哪裡?")) still seems to parse just fine. so most likely just using the passed input to parse.
- LikeAnElephant 6y agoYep, the browser has the UCD info built into it (a simplification... but basically). Similarly our mobile devices and various backend languages have the same data backed into it. This is where there are sometimes discrepancies between how a given browser or device would output this data, as it could be working off of an outdated version of Unicode's data. Some devices even overwrite the default Unicode behavior. There are just SO many languages and SO many regions and SO many combinations thereof that even Unicode can't cover all the bases. It's all very fascinating from an engineering perspective.
- yorwba 6y agoThe underlying library is actually using a single dictionary for both Chinese and Japanese https://github.com/unicode-org/icu/tree/7814980f51bca2000a963307cb5c4d711cc05fdb/icu4c/source/data/brkitr/dictionaries https://github.com/unicode-org/icu/tree/7814980f51bca2000a96...
- erjiang 6y agoIt turns out that's because ICU uses a combined Chinese/Japanese dictionary instead of separate dictionaries for each language. Which probably is a little more robust if you misdetect some Chinese text as Japanese and vice-versa.
- erjiang 6y agoThe library that Chrome uses seems to use a dictionary[0], since you can't determine word boundaries in Japanese just by looking at two characters. Your first link also says: > To handle certain situations, some line breaking implementations use techniques that cannot be expressed within the framework of the Unicode Line Breaking Algorithm. Examples include using dictionaries of words for languages that do not use spaces [0] posted in another top-level comment: http://userguide.icu-project.org/boundaryanalysis http://userguide.icu-project.org/boundaryanalysis
- LikeAnElephant 6y agoYea, CJK (Chinese, Japanese, Korean) breaking is particularly complex. Google has done a lot of work, and have this open source implementation which uses NLP. It's the best I've personally come across: https://github.com/google/budou https://github.com/google/budou
- oehtXRwMkIs 6y agoCan't imagine it being difficult for Korean.
- Asooka 6y agoKorean is especially difficult. Chinese uses only hanzi and you have a limited set of logogram combinations that result in a word. Japanese is easier because they often add kana to the ends of words (for e.g. verb conjugation) and you can use that (in addition to the Chinese algorithm) to delineate words - there is no word* where a kanji character follows a kana character. Korean on the other hand uses only phonetic characters with no semantic component, so you just have to guess the way a human guesses. * with a few exceptions, of course :)
- qiqitori 6y agoKorean uses spaces.
- 6y ago