3 ms·
Just a thought experiment, don't take it too seriously: The crux of the issue is that kanji don't have an inherent "natural" ordering that a user would expect.
by cooper12 7y ago
Just a thought experiment, don't take it too seriously:
The crux of the issue is that kanji don't have an inherent "natural" ordering that a user would expect. Sorting by their character code doesn't mean anything to a Japanese person. But, what if we made our own standard of what entails a "natural order". There's nothing about A–Z that makes the alphabet obligated to be in that order (and not something like based on sound or shape) other than it being the convention that developed. Even hiragana can have different orderings (AIUEO vs IROHA [0])
One proposed method would be to do it how the dictionaries do it: first sort by major radical, [1] and then by stroke count. This is something most Japanese learn when learning how to write characters anyway (of course ambiguities would arise when the radical is shared and the stroke count is the same, but we could just choose a third arbitrary factor; we'd also have to decide on a specific written form as stroke count can differ depending on whether it is handwritten or the font).
We could then teach our approach to schoolchildren and it would just become accepted over time like other things they learn. But wait you say, it's more natural for them to sort on pronunciation. However, if I gave you a list of polygon names and told you to sort by the number of sides they had, you'd be perfectly capable of doing it despite that not being alphabetical. Things are less "unnatural" if you grew up learning them and your brain doesn't experience dissonance.
Anyway, just my hot take.
[0]: https://en.wikipedia.org/wiki/Iroha https://en.wikipedia.org/wiki/Iroha
[1]: https://en.wikipedia.org/wiki/Radical_(Chinese_characters) https://en.wikipedia.org/wiki/Radical_(Chinese_characters)
- mjevans 7y agoThe concept of having a displayed value and a value to actually sort by isn't limited to Japanese words/names. It also comes up when there's a desired display function but for sorting intents different parts should be added or removed from the name of an entity. One such example is book and movie titles in a library.
- innocenat 7y ago> Even hiragana can have different orderings (AIUEO vs IROHA [0]) Unless otherwise mentioned, things are almost always AIUEO nowadays.
- jpatokal 7y agoThe only problem with that is that the resulting order would be useless for many applications. Say you're looking up your friend Tanaka Tarou, but you're not sure which characters his name is written with. If the sort order is phonetic, you can find the name and likely work out that this is the Tanaka you were looking for. But how do you search for a name in a kanji-indexed list if you don't know the kanji? Incidentally, this is why kanji dictionaries invariably have multiple indices: one by radical, the other by pronunciation(s).
- cooper12 7y agoGreat point. I was considering a very visual-minded reader but of course not everyone would be so good at it nor would they always just care about how the kanji looks rather than other aspects. It's a difficult problem indeed... My intention is to show that we'd need some sort of radical solution that might not be what we'd immediately jump to (for example the current approach the author mentions is having a separate field for readings, but this is clearly resource-intensive and wouldn't work on arbitrary data). To solve sorting for Japanese, I feel we need to rethink what it means to sort.