4 ms·
Ligatures, which are mentioned in the article, are a good example of the distinction between canonical and compatibility forms. If your input contains "traffic
by SloopJon 8y ago
Ligatures, which are mentioned in the article, are a good example of the distinction between canonical and compatibility forms.
If your input contains "traffic", you don't necessarily want NFC to insert the "ffi" ligature and turn it into "tra\ufb03c". The ligature is generally a poor choice in a monospaced font, for example.
Similarly, if you've gone to the trouble to insert the ligature, you don't necessarily want NFD to strip it out. However, with NFKC or NFKD, a search for "traffic" will find the string.
- lelf 8y agoLigatures are not a good example actually. They are obsolete by Unicode, i. e. there's no way to turn ffi⟨normal⟩ into ffi⟨ligature⟩. (But NFKD(ffi⟨ligarure⟩)≡ffi⟨normal⟩ of course.)
- jrochkind1 8y agoYou're right that there's no way to apply any specified unicode normalization to turn `ffi`(normal) into `ffi`(ligature). But they're a good example of the difference between "canonical" and "compatible" normalization anyway, in the other direction. Which is really the only direction that matters to illustrate the difference anyway. NFKD (or NFKC too I think?) will, as you say, turn the ligatures into their "individual" forms. It's different glyphs, and there is at that point no way to know which it was "originally", and no way to convert it in the other direction. The "compatibility" normalizations are "lossy". The "canonical" normalizations on the other hand are basically 'lossless' with regard to "glpyhs". Unless you actually _cared_ that it was represneted with a combining diacritic before, which in 99.9% of cases you don't. You should have exactly the same symbols on the screen after a 'canonical' normalization. And for any x, NFC(x) == NFC(NFD(NFC(x)). The compatibility normalizations are super useful, because as the person up there mentioned, you often want a search query for `ffi` to match on `ffi` (and vice versa). But they are intended to lose symbolic representation (ffi and ffi are now the same thing with no way to distinguish), where the canonical normalizations are not. I wouldn't say the `ffi` ligature is "obsoleted" in any way. People still use it all the time. We both just included it in our comments, unicode was happy to support that. Unicode is so happy for you to use it, that it provides compatibility normalization to make it _easier_ to recognize that it means the same thing as "ffi". :)