3 ms·
For those who want to know more about this: The article touches on just the tip of the iceberg. You might think that all you need to do is add an extra field f
by level3 7y ago
For those who want to know more about this:
The article touches on just the tip of the iceberg. You might think that all you need to do is add an extra field for phonetic readings, and then simply sort on that field, but there are a lot of things that can go wrong. A naive sort (i.e. based simply on character code) will hit the following snags:
1) Hiragana vs Katakana
The article focuses more on Kanji vs Kana, but Japanese users will expect Hiragana and Katakana to be properly sorted together. Either you normalize your sort field (by converting everything to Hiragana, for example) or you use a Kana-insensitive collation.
2) Half-width characters
Katakana can be encoded as full-width or half-width characters (カ vs カ). Generally you want these treated as the same, so again you need to normalize or use a width-insensitive collation. There are also full-width alphabet characters (A vs A).
3) Youon
These are actual different characters (ゆ vs ゅ, つ vs っ), so you can't normalize, but you want them sorted together. Here you need a collation that's case-insensitive with respect to these.
4) Dakuten/Handakuten
Like youon, these are also different characters (は vs ば vs ぱ) so you can't normalize, but you want them sorted together (insensitively). A sensitive sort will give you (はね, ばつ, ぱすた) while an insensitive sort will give you (ぱすた, ばつ, はね).
There has been a lot of work around this, resulting in many different database collations over the years, some of which result in sorts that would greatly confuse Japanese users. As of today, you probably want to be using (in the case of MySQL) utf8mb4_0900_ai_ci or something similar.
- txtsd 7y agoI'd assume you'd want ゅ andっ to be added to the kana they're attached to, and then sorted. I'd want my き and きゅ together and び and っび together.
- level3 7y agoThat does make sense in a way, but I don't think that would feel natural to any native Japanese speaker. I'm not native and even I would find that ordering very odd. At the very least, it would make the sorting algorithm a lot more complex if you had to look ahead at later characters in order to sort the current prefix.