5 ms·
> but accent-insensitive really makes sense for (probably) most applications, considering that's also how dictionaries are sorted in german. I doubt it's exact
by chimeracoder 4y ago
> but accent-insensitive really makes sense for (probably) most applications, considering that's also how dictionaries are sorted in german.
I doubt it's exactly insensitive in the way that's described in the article, because it's still likely deterministic, whereas in the article the behavior can be nondeterministic.
I'm having a trouble thinking of an example in German, but in Spanish, there are plenty of words that differ only by an accent mark: "papá" and "papa", "tú" and "tu", etc. I'm guessing that any given dictionary would still treat these identical-but-for-diacritic words consistently throughout the dictionary, rather than sometimes placing the one with the umlaut first and sometimes placing the one with the umlaut second.
- theamk 4y agoThe article's behavior is non-determenistic only because author didn't finish designing the system and just kinda gave up instead. For example in the last statement, one approach might be to keep all equivalent strings, then sort them using some other complete ordering function (like utf-8 codepoints) and always return first one. (Yes, this will increase complexity, but locales often do that)
- Sharlin 4y agoBut if you use eg. GROUP BY as in the author’s example it’s not really up to you which representative of the equivalence class is chosen (I wonder if there’s a Unicode algorithm for transforming a word to a canonical representative of its EC… probably not in the general case).
- theamk 4y agoIf you are a database user, sure, you are at mercy of database developer. But if you are writing a database yourself, it is entirely up to you which representative will be returned by GROUP BY. You can take an easy way out and implement "return non-deterministic random member of the group" strategy. Or you can choose the representative with smallest UTF-8 point values, so it will be fully deterministic (I'd prefer this one personally). Or choose row with smallest insertion date. Or let user choose. There are tons of options.
- Dylan16807 4y agoIs a dictionary ever going to be sorting those words against each other outside of a single copy in the main list? You're always consistent if you only do it once.
- chimeracoder 4y ago> Is a dictionary ever going to be sorting those words against each other outside of a single copy in the main list? You're always consistent if you only do it once. If a dictionary listed "papa" before "papá" but "más" before "mas", I would call that an inconsistent lexicographic ordering.
- thematrixturtle 4y agoCompound words are quite common: Papá Noel, papaíto, papas fritas, papista
- Dylan16807 4y agoIf you're counting words with suffixes, I already checked the oxford spanish-english dictionary as the first useful google result, and it goes back and forth between accent and no accent, for example papa then papá then papada. But that's not really the same thing as having the entire words papa and papá be sorted one way in one place, and another way in another place.
- tragomaskhalos 4y agoThe acute in Spanish is a stress marker is it not, therefore (to this non-expert anyway) it makes more sense to treat the accented and unaccented vowels the same for collation purposes than would be the case for languages where the diacritic denotes a different phoneme. That said, I don't know enough Spanish to know whether the acute ever distinguishes separate words, which would perhaps weaken that argument. The grave in Italian serves the same purpose AFAIK so would be interesting to compare the rules for that language.
- 4y ago
- ygra 4y agoUmlauts are always sorted like the base letter + e in German, so as are, or, ue. However, I think with accent-imsensitive they meant that accents that don't exist in German are simply ignored when sorting. So you wouldn't have German words where the letters only differ by accent. Umlauts are normal German letters, so they have sorting rules (albeit different ones from the same letter in other countries, e.g. in Swedish I think they simply come at the end of the alphabet).
- lokedhs 4y agoYou're correct about Swedish. The extra letters come at the end, in the order å, ä, ö. Interestingly enough this order is different to how Norwegian and Danish orders them. Also, as opposed to German, they are not referred to as umlauts or accents. ä is a completely different letter compared to a, and they have absolutely no relation. An indication of this is that I can't even think of a word that describes the dots over the a, similar to how most English speakers wouldn't be be able to name the dot over the i. As such, a search or sorting algorithm that puts equivalence between ä and a would be completely broken. However, an English speaker who only wants to look up a Swedish word in a dictionary would definitely want that equivalence. They may not even be able to type the Swedish letters. From this I conclude that it's the locale of the user that must dictate the way comparison works, not the language of the words that are being compered. This is thankfully how locales typically work. In Swedish, these are extremely common letters, there are plenty of examples where words differ only in letters which would change using such algorithms. Such as älg/alg (elk/algae), or kö/ko (queue/cow).
- rags2riches 4y agoIt's not exactly right to say that a/å/ä and o/ö have absolutely no relation in Swedish. There are plenty of language cases where they act more like umlauts, such as stor/större (large/larger) or få/färre (few/fewer). Despite this, they are still always treated as separate letters. Also, the history of the letter shapes are the same as the german umlauts, with e being written above a or o before writing ä or ö and o written above a before writing å.
- jcranmer 4y agoFrench happens to have a quartet of words that makes collation order obvious: cote, coté, côte, and côté. (That's the order my French-English dictionary orders those words; my French dictionary swaps coté and côte.)