4 ms·
Yes. This seems to work even when you pass it Chinese while maintaining ja-JP as the language function tokenizeJA(text) { var it = Intl.v8BreakIterator([
by swang 6y ago
Yes. This seems to work even when you pass it Chinese while maintaining ja-JP as the language
function tokenizeJA(text) {
var it = Intl.v8BreakIterator(['ja-JP'], {type:'word'})
it.adoptText(text)
var words = []
var cur = 0, prev = 0
while (cur < text.length) {
prev = cur
cur = it.next()
words.push(text.substring(prev, cur))
}
return words
}
console.log(tokenizeJA("今天要去哪裡?"))
still seems to parse just fine. so most likely just using the passed input to parse.
- LikeAnElephant 6y agoYep, the browser has the UCD info built into it (a simplification... but basically). Similarly our mobile devices and various backend languages have the same data backed into it. This is where there are sometimes discrepancies between how a given browser or device would output this data, as it could be working off of an outdated version of Unicode's data. Some devices even overwrite the default Unicode behavior. There are just SO many languages and SO many regions and SO many combinations thereof that even Unicode can't cover all the bases. It's all very fascinating from an engineering perspective.
- yorwba 6y agoThe underlying library is actually using a single dictionary for both Chinese and Japanese https://github.com/unicode-org/icu/tree/7814980f51bca2000a963307cb5c4d711cc05fdb/icu4c/source/data/brkitr/dictionaries https://github.com/unicode-org/icu/tree/7814980f51bca2000a96...
- erjiang 6y agoIt turns out that's because ICU uses a combined Chinese/Japanese dictionary instead of separate dictionaries for each language. Which probably is a little more robust if you misdetect some Chinese text as Japanese and vice-versa.