4 ms·
How is defaulting to byte-comparing strings Anglocentrist?
by Denvercoder9 5y ago
How is defaulting to byte-comparing strings Anglocentrist?
- viraptor 5y agoThe idea that you can byte-compare strings is ascii-centrist. Others will ask "are they even using the same encoding", "are they normalised the same way" first.
- Denvercoder9 5y agoThose things only matter when your strings contain human-oriented text, but that's not the only (and dare I say not even most common) usage of strings. Obviously it's best to be explicit about how you want the contents of your strings to be treated, but absent that, I don't see how defaulting to ordinal string comparison is Anglo-centric. The English culture doesn't use ordinal comparison. Defaulting to ordinal comparison can actually _prevent_ breakage for people using non-English cultures, because strings suddenly behave differently in ways the developer didn't anticipate or test. (Also, I should've said codeunit-compare instead of byte-compare, since C# strings are always UTF-16 encoded).
- jameshart 5y ago.NET strings are backed by two-byte UTF-16 codepoints, not bytes in an unknown encoding. But ordinal comparison means ignoring equivalence classes, like U+006E (n) + U+0303 (combining tilde) being considered the same as U+00F1 (ñ). But - perhaps more surprisingly - .net's non-ordinal match also matches things like "ß" and "ss". That leads to some weird things like even if two strings compare as equal in the invariant culture, if you uppercase them in the invariant culture they may no longer be considered equal. And that's all before you start looking at culture-variant behaviors like different cultures having different casing rules (see Turkish dotless i) or approaches to diacritic equivalence for sorting.