4 ms·
In Hungarian "dzs", "zs" and "sz" are so called compund letters, however most software treats them as separate characters, because it would be a pain in the ass
by oldsecondhand 5y ago
In Hungarian "dzs", "zs" and "sz" are so called compund letters, however most software treats them as separate characters, because it would be a pain in the ass do sorting the correct way. (It would either require changing the input method, or software would need to be aware of etymology.) It's only sorted the right way in paper dictionaries.
- dhosek 5y agoActually, it just needs to be context sensitive. If I have an English document and I have an index which has an entry for Zsigmond, it would be placed between Zebra and Zygote. But, if I have a Hungarian document, then the correct ordering would be to have Zsigmond after Zebra and Zygote. Any correctly-implemented program that deals with sorting should not be assuming it can just sort on code point¹ and Hungarian is one of the standard collations that is part of the Unicode reference implementations so it shouldn't be a problem. The fact that a lot of software is badly-implemented is not an excuse for software to continue to be so. I did a quick check and saw that I can specify Hungarian in MySQL as a collation for a table and I also know that it's available in Java and even JavaScript. ⸻ 1. This sorting is wrong not just for Hungarian but for every language including English which would expect, e.g., naïve to be sorted between nag and nanny and not after nay.