4 ms·
I'm... honestly not sure. I think I got the right word, but I don't know a breath of Japanese, so I might have copied the wrong one. At any rate, the results
by JasonFruit 4y ago
I'm... honestly not sure. I think I got the right word, but I don't know a breath of Japanese, so I might have copied the wrong one. At any rate, the results of trying it just now were seriously underwhelming.
- derefr 4y agoThe trouble with Japanese in particular, but ideograph languages generally, is that the lack of explicit spaces makes tokenization nontrivial (you essentially need to parse words out using a model built on a dictionary) — and English-language websites don’t even bother to think about doing this when trying to build their search engines; they just throw things into Lucene or Postgres tsvector and think the problem solved. This results in indexing document titles in these languages as if they were bags of individual characters (i.e. splitting on every character.) Other ideograph languages (e.g. Chinese) aren’t so bad when you do this, because enough information is captured per ideograph that losing ordering doesn’t actually change meaning that much. But Japanese is especially bad, because it often “reverts” to the hiragana or katakana alphabets for long strings of words, while still not putting any explicit word-break markers between saidwords. So you end up indexing a “bag of letters.” You can imagine the uselessness of such results. Or, to put that another way: Japanese-language search “works” on English-language sites when you’re searching a term expressed in kanji. But when the term you want turns out to only have a katakana representation, it’s mostly not going to work. And yes, Google themselves know all this and have solved all these problems for Google Search long ago. Apparently, the relevant tech never made it into YouTube.