5 ms·
In some cases pretending everything is ASCII is the sane thing to do. With Unicode, sorting and case conversion are neigh impossible to do correctly. While ther
by rustybolt 2y ago
In some cases pretending everything is ASCII is the sane thing to do. With Unicode, sorting and case conversion are neigh impossible to do correctly. While there are algorithms for collating codepoints into (extended) grapheme clusters, there is still a lot of freedom, so while there are wrong ways to do it there is no canonical right way.
- poincaredisk 2y ago>With Unicode, sorting and case conversion are neigh impossible to do correctly Surely you mean that sorting correctly is impossible without Unicode? Otherwise you would have to hardcode the rules of sorting strings correctly in my language (and all other languages) yourself. Unless your preferred solution is "close my eyes and prefer non-ascii characters don't exist", then... I'm not a fan.
- samatman 2y agoSorting is impossible to do correctly without knowledge of the language in which the text is written, because the collation rules for symbols differ between languages. Unicode, of course, defines those collation rules, and UTF-8 sorts lexicographically using the same naïve byte comparison which works for ASCII. Case conversion is similar except the default rules do a very good job in general. But still, there are a few language-specific quirks and, again, you do have to know what language is involved to get those right. I'm agreeing with you, to be clear, just adding that a) Unicode isn't always enough, but it does a decent job if you don't know the language in advance, and that it defines the correct rules if you do know that.
- sgarland 2y ago> UTF-8 sorts lexicographically using the same naïve byte comparison which works for ASCII This isn’t necessarily true beyond ASCII, and it depends entirely on the collation [0]. One need only to spend some time peering into the abyss that is RDBMS collation support [1] [2] to see the horror. [0]: http://www.unicode.org/reports/tr10/ http://www.unicode.org/reports/tr10/ [1]: https://dev.mysql.com/doc/refman/8.4/en/charset-unicode-sets.html https://dev.mysql.com/doc/refman/8.4/en/charset-unicode-sets... [2]: https://www.postgresql.org/docs/current/collation.html https://www.postgresql.org/docs/current/collation.html
- samatman 2y agoWell, no. It sorts according to the lexiographical order of Unicode. Earlier points before later points. How useful this is depends on the language, of course. I did say that as well. But Unicode was put together from legacy character encodings, and did what it could to preserve the order of those, so it's far from useless.
- teddyh 2y ago“Sorting is hard in other languages, so I would like to force everybody to only use the characters from my language, no characters from any other language. This will make it easy for me.”
- Etherlord87 2y agoEnglish is the international language, and latin characters have a special 'canonical' status in computer science, so you're heavily strawmaning here...
- DrillShopper 2y agoLaziness, Impatience, and Hubris