4 ms·
I was really surprised when realized that at least in hpfs cyrillics is normalized too. For example, no russian ever thinks that Й is a И with some diacritics.
by codesnik 3y ago
I was really surprised when realized that at least in hpfs cyrillics is normalized too. For example, no russian ever thinks that Й is a И with some diacritics. It's a different letter on it's own right. But mac normalizes it into two codepoints.
- anamexis 3y agoWell, there's no expectation in unicode that something viewed as a letter in its own right should use a single codepoint.
- asveikau 3y agoI dislike explaining string compares to monolingual English speakers who are programmers. Similar to this phenomenon of Й/И is people who think ñ and n should compare equally, or ç and c, or that the lowercase of I is always i (or that case conversion is locale-independent). In something like a code review, people will think you're insane for pointing out that this type of assumption might not hold. Actually, come to think of it, explaining localization bugs at all is a tough task in general.
- iforgotpassword 3y agoWell, I do like this behavior for search though. I don't want to install a new keyboard layout just to be able to search for a Spanish word.
- david-gpu 3y agoIs the convenience of a few foreigners searching for something more important than the convenience of the many native speakers searching for the same? Maybe we should start modifying the search behavior of English words to make them more convenient for non-native speakers as well. We could start by making "bed aidia" match "bad idea", since both sound similar to my foreign ears.
- MrJohz 3y agoIn fairness, for search, allowing multiple ways of typing the same thing is probably the best choice: you can prioritise true matches, where the user has typed the correct form of the letter, but also allow for more visual based matches. (Correcting common typos is also very convenient even for native speakers of a language — and of course a phonetic search that actually produced good results would be wonderful, albeit I suspect practically very difficult given just how many ways of writing a given pronunciation there might be!)
- david-gpu 3y agoAs a counterexample, conflating two different glyphs as if they were the same can lead to the inability to search for a particular term. E.g. in Spanish these two words (cono, coño) have very different meanings. If I'm searching for one I don't want to see results pertaining to the other one. It would be like searching for "sheet" and getting results for "shit".
- MrJohz 3y agoIt depends on how the search is implemented exactly and what the context is, but assuming I've searched for "cono", I would expect results that directly match "cono" to come first, then results that also match "coño". Similarly to how I'd expect to still get reasonable results if I type "beleive" instead of "believe". That said, this is obviously pretty context-dependent, in some settings it will make more sense to do an exact-match search, in which case you'd want to differentiate n and ñ (while still handling different possible unicode variants of ñ if those exist).
- makeitdouble 3y agoSearch probably needs both modes. A literal and a fuzzy one.
- dmckeon 3y agoFor similar sounding names, this fuzzy match is pretty effective. https://www.archives.gov/research/census/soundex https://www.archives.gov/research/census/soundex
- nradov 3y agoIn terms of phonetic matching algorithms, Soundex is considered badly outdated. Most MDM products use more advanced alternatives.
- NeoTar 3y agoMy brother recently asked for help in determining who a footballer (soccer player) was from a photo. Like in many sports, the jerseys have the players name on the rear, and this player’s was in Cyrillic - Шунин (Anton Shunin) - and my brother had tried searching for Wyhnh without success. Anyway, my point is that perhaps ideally (and maybe search engines do this) the results should be determined by the locale of the searcher. So someone in the English speaking world can find Łódź by searching for Lodz, but a Pole may need to type Łódź. My brother could find Shunin by typing Wyhnh, but a Russian could not…
- nradov 3y agoEssentially you are asking for search engines to recognize "Volapuk" encoding. https://en.wikipedia.org/wiki/Informal_romanizations_of_Cyrillic#%3A%7E%3Atext%3DVolapuk_encoding_%28Russian%3A_%D0%BA%D0%BE%D0%B4%D0%B8%D1%80%D0%BE%D0%B2%D0%BA%D0%B0_%22%2CCyrillic_script_with_Latin_ones.?wprov=sfla1 https://en.wikipedia.org/wiki/Informal_romanizations_of_Cyri...
- yxhuvud 3y agoOr that sort order is locale independent. Swedish is a good example here as åäö are sorted at the end, and where until 2006 w was sorted as v. And then it changed and w is now considered a letter of its own.
- makeitdouble 3y agoThe general reaction I've see until now was "meh, we have to make compromises (don't make me rewrite this for people I'll probably never meet)" Diacritics exacerbate this so much as they can be shared between two language yet have different rules/handling. French typically has a decent amount and they're meaningful but traditionally ignores them for comparison (in the dictionary for instance). That makes it more difficult for a dev to have an intuitive feeling of where it matters and where it doesn't.
- koliber 3y agoThese are different letters for people who speak the language and treating them the same in some usage seems weird. At the same time, sometimes words containing those letters might show up in context where the user is not familiar with that language. Such users might not know how to enter those letters. They might not even have the capability to type those letters with their installed keyboard layouts. If they are searching for content that contains such letters (e.g. a first name), normalizing them to the visually-closest ASCII is a sensible choice, even if it makes no sense to the speakers of the language. It's important to understand a situation from different perspectives. It's not about coming up with a single correct interpretation that makes logical sense. It about making a system work in least-surprising ways to all classes of users.
- bawolff 3y agoNormalization isn't based on what language the text is. NFC just means never use combining characters if possible, and NFD means always use combining characters if possible. It has nothing to do with whether something is a "real" letter in a specific language or not. The whether or not something is a "real" letter vs a letter with a modifier, more comes into play in the unicode collation algorithm, which is a separate thing.