14 ms·
Unicode Normalization Forms: When ö ≠ ö
- guerrilla 5y ago> But here, normalization caused this issue. Nope, the lack of normalization on both accounts by the SMB server caused the issue. It could have normalized before emitting but it definitely should have normalized on receiving for comparison.
- silon42 5y agoAt least it should perform validation and reject the NFD form and force the client to normalize to NFC?
- rurban 5y agoNFD is fine when you dont have much time and can afford the space. NFC is about 3x slower and smaller. Forcing clients never works. Be tolerant what you accept and strict what you write. In the case of C++23 enforcing NFC my mind is twisted. It would allow heavy tokenizer optimizations, but this is an offline compiler, where you don't really need that. The problem are the compatible variants, NFKC and NFKD. But then you have usecases where you need them, and more to actually find strings. Levenshtein should not be the default when searching for strings.
- B-Con 5y agoI think that in the ls->read workflow, Nextcloud shouldn't normalize the response from SMB and should issue back to SMB whatever SMB returned to Nextcloud.
- guerrilla 5y agoAccording to Unicode, it should be allowed to and the SMB server should be able to handle it. That's kind of the point of normalization, they're meant to be done before all comparisons so that exactly this doesn't happen. Your suggestion is just premature optimization, i.e. eliminating a redundancy.
- int_19h 5y agoUnicode doesn't say anything about what "should be allowed to" with respect to an unrelated protocol. If the protocol says that filenames are sequences of 16-bit values that have to be compared one by one, then that's what it is.
- guerrilla 5y agoIt does say that if comparisons are being made then... and comparisons are being made, so yes, it does.
- int_19h 5y agoIf comparisons are being made of Unicode strings, sure. Does the protocol actually defines the identifier in question as a Unicode string, though? Or as an array of 16-bit ints?
- misnome 5y agoWhy isn’t the answer just “Don’t unicode normalise the file name”? I thought the generally recommended way to deal with file names is to treat as a block of bytes (to the extent that e.g. rust has an entirely separate string type for OS provided strings), or just to allow direct encoding/decoding but not normalisation or alteration.
- pavlov 5y agoThe most common desktop file systems are case-insensitive, which complicates the picture.
- Pxtl 5y agoStill, it looks like the right thing to do is let the filesystem do the filesystem's job. The filesystem should be normalizing unicode and enforceing the case-insensitivity and whatnot, but just the filesystem. Wrappers around it like whatever Nextcloud is doing should be treating the filenames as a dumb pile of bytes.
- dataflow 5y agoI'm not sure this problem even has a "right" solution. > Wrappers around it like whatever Nextcloud is doing should be treating the filenames as a dumb pile of bytes. What do you do when the input isn't a dumb pile of bytes, but actual text? (Like from a text box the user typed into?)
- nixpulvis 5y agoHalf normal isn't normal. That said, I personally try to avoid unicode in filenames (and caps too) for similar reasons.
- 0x0 5y agoJava is terrible in this regard, as most file APIs use "java.lang.String" to identify the filename, which most of the time depends on the system property "file.encoding". With the result that there will be files that you can never read from a java application if the filename encoding does not match the java file.encoding encoding.
- mgaunard 5y agoMost formats (including XML) require data to be normalized to NFC.
- chrismorgan 5y agoCan you point me to a single format that actually requires NFC? Most things either make no comment or just express preferences, though I’m confident there will be some somewhere. XML does not require normalisation: per <https://www.w3.org/TR/xml11/#sec-normalization-checking https://www.w3.org/TR/xml11/#sec-normalization-checking>, XML data SHOULD be fully normalised, but MUST NOT be transformed by processors; in other words, it’s a dead letter “SHOULD”, and no one actually cares, just like almost everything else.
- heikkilevanto 5y agoWell, if 7-bit US ASCII was good enough for our Lord, it is good enough for me ;-)
- kingcharles 5y agoWell.. if we're getting technical, the "Old" Testament is written in Hebrew, and the "New" Testament in written in Greek. The first line of Genesis reads thus: (from right-to-left, although the earliest Hebrew is actually LTR) בְּרֵאשִׁית, בָּרָא אֱלֹהִים, אֵת הַשָּׁמַיִם, וְאֵת הָאָרֶץ. And the beginning of the Gospel of Mark thus: Ἀρχὴ τοῦ εὐαγγελίου Ἰησοῦ Χριστοῦ υἱοῦ θεοῦ. (and if we're getting super technical there are a bunch of Aramaic phrases in the Bible that Jesus spoke, although I know little of Aramaic and I don't know how it would have been written in Biblical times between the Greek characters) So the Lord would be needing those 16-bits after all..
- javajosh 5y agotl;dr - don't use crazy unicode characters in filenames, they can be problematic for non-trivial reasons (in this case because of unicode normalization on an smb mount.)
- int_19h 5y agoWhat's "crazy" about the letter? It's a standard letter of several European alphabets.
- drpixie 5y agoNothing crazy about the "letter", but it is crazy that there are multiple different ways to encode the "letter".
- dagmx 5y agoA user wouldn't know that there are multiple ways to encode a given character unless they're experienced with Unicode. Additionally there are (iirc) multiple ways to encode characters even in the ASCII set. This is purely a failing in consistent normalization schemes.
- kortex 5y agoSo, no combining characters? Ok, even if you rule out Latin characters with accents and only use code points that consider the character with accent a single entity... you still have world languages which need combining in order to work, which means you can't really escape multiple encodings of the same "graphemes" (these languages don't exactly have "letters" like ASCII does).
- hinkley 5y agoReading about unicode has made me much, much more circumspect about the meaning of != in languages, and what fall-through behavior should look like. Unicode domain names lasted for a hot minute until someone registered microsoft.com with Cyrillic letters. Years ago I read a rant by someone who insisted that being able to mix arbitrary languages into a single String object makes sense for linguists but for most of us we would be better off being able to a assert that a piece of text was German, or sanskrit, not a jumble of both. It's been living rent free in my head for almost two decades and I can't agree with it, nor can I laugh it off. It might have been better if the 'code pages' idea was refined instead of eliminated (that is, the string uses one or more code pages, not the process). I don't know what the right answer is, but I know Every X is a Y almost always gets us into trouble.
- int_19h 5y agoYou can already map Unicode ranges to "code pages" of sorts, so how would that help? Thing is, people who are not linguists do want to mix languages. It's very common in some cultures to intersperse the native language with English. But even if not, if the language in question uses a non-Latin alphabet, there are often bits and pieces of data that have to be written down in Latin. So that "most of us" perspective is really "most of us in US and Western Europe", at best. For domains and such, what I think is really needed is a new definition of string equality that boils down to "are people likely to consider these two the same?". So that would e.g. treat similarly-shaped Latin/Greek/Cyrillic letters the same.
- jrochkind1 5y agoOh, you can do far more than "code pages of sorts". Unicode has a variety of metadata available about each codepoint. The things that are "code pages of sorts" are maybe "block" (for ö "Latin-1 Supplement"), and "plane" (for ö it's "Basic Multilingual Plane"), but those are really mostly administrative and probably not what want. But you also have "Script" (for ö "Latin). Some characters belong to more than one script though. Unicode will tell you that. Unicode also has a variety of algorithms available already written. One of the most relevant ones here is... normalization. To compare two strings in the broadest semantic sense of "are people likely to consider these the same", you want want a "compatibility" normalization. NFKC or NFKD. They will for instance make `1` and `¹`[superscript] the same, which is definitely one kind of "consider these the same" -- very useful for, say, a search index. That won't be iron-clad, but that will be better than trying to role your own algorithm involving looking at character metadata yourself! But it won't get you past intentional attacks using "look-alike" characters that are actually different semantically but look similar/indistinguishable depending on font. The trick is "consider these the same" really, it turns out, depends on context and purpose, it's not always the same. Unicode also has a variety of useful guides as part of the standard, including the guide to normalization https://unicode.org/reports/tr15/ https://unicode.org/reports/tr15/ and some guides related to security (such as https://unicode.org/reports/tr36/ https://unicode.org/reports/tr36/ and http://unicode.org/reports/tr39/ http://unicode.org/reports/tr39/), all of which are relevant to this concern, and suggest approaches and algorithms. Unicode has a LOT of very clever stuff in it to handle the inherently complicated problem of dealing with the entire universe of global languages that Unicode makes possible. It pays to spend some time with em.
- jrochkind1 5y agoit seems like a bug that to get consistent unicode normalization you need to flip a non-default config option. What am I missing?
- mannerheim 5y agoDuolingo doesn't handle Unicode normalisation for certain languages, and it's incredibly frustrating. Here's one example[0] (Vietnamese) and I know it's the case for Yiddish as well. [0]: https://forum.duolingo.com/comment/17787660/Bug-Correct-Vietnamese-input-and-no-Unicode-normalization-gives-incorrect-response https://forum.duolingo.com/comment/17787660/Bug-Correct-Viet...
- tpmx 5y agoAs a northern European, I kinda miss iso-8859-1 being used everywhere back in the mid 90s.
- sharikous 5y agoAs a Middle-Eastern I still dread those times
- chrismorgan 5y agoA fun related issue that could occur: applying NFD to a string can make it longer, so a sanitiser that limits file names to 255 UTF-16 code units but doesn’t first normalise to NFD could fail on HFS+. This could occur on systems that normalise to NFC as well: NFC lengthens some strings, e.g. 𝅗𝅥 (U+1D15E MUSICAL SYMBOL HALF NOTE) normalises to 𝅗𝅥 (U+1D157 MUSICAL SYMBOL VOID NOTEHEAD, U+1D165 MUSICAL SYMBOL COMBINING STEM) in both NFC and NFD (similar happens in various Indic scripts, pointed Hebrew, and the isolated case of U+2ADC FORKING which is special for reasons UAX #15 explains), but I don’t think there are any file systems that actually normalise to NFC? (APFS prefers NFC, but doesn’t normalise at the file system level.) The remaining concern would be that NFC could take more UTF-8 code units than NFD despite adding a character, but in practice this doesn’t occur (checked on NormalizationTest-3.2.0.txt).
- kingcharles 5y agoThat's a really great point about the string-length and not often addressed. You might even be able to force some sort of buffer overflow with that I guess.
- vvhn 5y agoAPFS doesn’t “prefer” anything - it it will not change the bytes passed to NFC or NFD. The bytes passed for creation are stored as is ( HFS will store the NFD form on disk if you pass the NFC form to it). However APFS is normalization insensitive (if you create a NFC name on disk , you won’t be able to create the NFD version and you will be able to the name by both the NFC and NFD variants) just as HFS is - they both use different mechanisms to achieve normalization insensitivity.
- rurban 5y agoA filesystem accepting only NFD should be filed as bug. They can normalize it internally to NFD, as Apples previous HFS+ did. But even worse than that is Python's NFKC, which normalizes ℌ to H and so on. The recommended normalizations are NFC for offline normalization (like in compiled languages and databases) and NFD for online, where speed trump's space. unicode.org talking that much about NFKC was a big mistake. NFKC is crazy and doesn't even roundtrip. The whole TR31 XID_Start/Continue sets are mostly because of NFKC issues, not so about stability. But people bought it for its stability argument. I'm just writing a library and linter for such issues: https://github.com/rurban/libu8ident https://github.com/rurban/libu8ident Also note that C++23 will most likely enforce NFC identifiers only. Same problem as with this filesystem. My implementation was to accept all normal. forms and store it internally and in the object files as NFC. The C ABI should declare it also. Currently they don't care as much as Linux filesystems: Nada. Identifiers being unidentifiable
- WalterBright 5y agoWhen Unicode adopted normalization it went off the rails into mudville. Then, determined to make a mockery of its purpose, it adopted semantic meanings, fonts, and then started inventing all sorts of new characters. "If you vote for my nutburger glyph my kid drew for a kindergarten assignment, I'll vote for the chicken scratching you noticed in the barnyard dust."
- a1371 5y agoEventually someone will write a string matching library that renders the characters on an internal canvas and diffs the pixels instead.
- tsimionescu 5y agoA major problem with Unicode is that it gives you strange ideas about text: that you can somehow take human text encoded in Unicode and answer questions like "how many letters does this have" or "are these two pieces of text different" or "split this text into words" in a way that works generically for any langauge or context. These are all myths, and APIs for such things are bugs. The only thing you can meaningfully do with two pieces of arbitrary Unicode text is to say if they are byte-by-byte equal. For any other operation, you need to have specific business logic. For example, are "Ionuț" and "Ionut" and "Ionutz" the same string or different strings? There is no generic answer: depending on the intended business logic, they may be identical or not (e.g. if we consider these to be Romanian names, they should be considered identical for search purposes, but probably not identical for storage purposes, where you want to remember exactly how the person spelled their name). A related problem is that most langauges have no separate types for Text/String on one hand, and Symbol on the other. Text or other Strings should be opaque human text, that can only be interpreted by specific code, offering almost no API (only code point iteration). Symbols should be a restricted subset of Unicode that can offer fuller features, such as lengths, equality, separation into words etc. This would be the type of JSON key names used in serialization and deserialization, for example.
- rstuart4133 5y agoWhat they should have done is not that strange. Text is merely an ordered collection of characters. If you just assigned each character (aka grapheme) a number, text becomes a sequence of numbers. The first two questions you pose, "how many letters does this have" and "are these two pieces of text different" are trivially answered by such a representation. Unicode's fuck up is the managed to come up with something that can not reliably answer those two questions. In fact what Unicode has end up with is so horrible, it's a major exercise in coding just to answer a simple question like "is there an 'o' in this sentence", as in Python3's "'o' in sentence" does not always return the right result. Unicode's starting point was all wrong. There is an encoding that did a perfectly good job of mapping graphemes to numbers: ISO-10646. In fact Unicode is based on it, by then committed their original sin: they decided all the proposed ISO-10646 encodings (ie, how the numbers are encoding into byte streams) were crap, so they released a standard that combined two concepts that should have remained orthogonal: codepoints and encoding those codepoints to a binary stream. Now it's true ISO-10646 proposed encodings were undercooked. That became painfully apparent when Ken Thompson came up with utf-8. But no biggie right: utf-8 was just another ISO-10646 encoding, just let it take over naturally. The Unicode solution to the encoding problem was to first decide we would never need more than 2^16 codepoints, then wrap it up in "one true encoding everyone can use": UCS2. Windows and Java, among others, bought the concept, and have paid the price ever since. They were wrong of course. 2^16 was not enough. So they replaced the USC2 encoding with UTF-16 which was sort of backwards compatible. But not one UTF-16, oh no, that would be too simple. We got UFT-16LE and UTF-16BE. Notice what has happened here: take identical pieces of text, encode them as valid Unicode, and end up with two binary objects that were different. Way to go boys! But that wasn't the worst of it: they managed to screw up UTF-16 so badly it didn't expand the code space to the 2^32 points, just 2^20. And in case you can't guess what happens next, I tell you: turns out there are more than 2^20 grapheme's out there. What to do? Well there are a lot of characters that are "minor variants” of each other, like, like o and ö. Now Unicode already had a single code point for ö but to make it all fit and be uniform they decided "Combining Diaeresis” was the way these things should be done in future. So now the correct way to represent ö is a code point that says "add an umlaut to the next character (provided it isn't another diaeresis)" followed by the code point for o. But as the original codepoint for ö still exists, we can have two identical graphemes that don't compare as equal under Unicode, which is how we get to ö ≠ ö. So it's not only Python3 "'o' in sentence" that doesn't always always work. We arrived at the point that "'ö' in sentence" can't be done without some heavy lifting that must be done by a library. Just to make it plain: some CPU's can do "'o' in sentence" in a single instruction. That simple design decison have lost is orders of magnitude in CPU efficiency. I know these are strong words, but IMO this is a brain dead, monumental fuckup, making things acre feet, furlong fortnights look positively sane. It's time to abandon Unicode, and it's “Combining Diaeresis” in particular and go back to basics: ISO-10646 and utf-8. UTF-8 provides a 28 bit encoding space, which is more than enough to realise the one the single guiding principle that ISO-10646 was founded on: one codepoint per grapheme. It won’t happen of course, so as a programmer I’ll have to deal with the shit sandwich the Unicode consortium has served up for the rest of my life.
- kingcharles 5y agoAnd then some jerk comes along and writes ꙮ one time, in one document, in the whole of human history, just for the lulz, and the next thing you know we're saying hello to a new log4j. https://en.wikipedia.org/wiki/Multiocular_O https://en.wikipedia.org/wiki/Multiocular_O
- h0nd 5y agoReminds me of the old pаypal.com scam, where а != a.