3 ms·
Couldn't you solve that by sorting by the underlying byte order? Since it is required to be UTF-8 that should be fast and unambiguous.
by amock 10y ago
Couldn't you solve that by sorting by the underlying byte order? Since it is required to be UTF-8 that should be fast and unambiguous.
- dsp1234 10y agoYou could, and that's partially what prompted the question, but that's not quite what the specification says currently. edit: and unambiguous. var x = "\u0307\u0323\u0073" var y = "\u0073\u0323\u0307" These are arguably the same grapheme when displayed on a screen, but they would be considered different keys in SON, and sorted in different spots. I'm actually unsure if most diffing systems would catch that they are different or not, and if so, how they would represent that difference to the end user (since just showing the grapheme would be useless as they are the same). The specification is silent on normalization, so a SON pipe could output an object with those as keys (and still meet spec), but not be accepted by a conforming SON parser (that did unicode normalization on incoming keys). So really, the specification should say which of the canonical normalized forms to use, or that keys should specifically may not be normalized/modified in any way. This works fine if you're running something like JavaScript -> JavaScript -> JavaScript, but would fail if a (currently conforming) system in the middle defaults to normalized unicode representations. In other words, would a pipeline like Firefox -> Swift -> Python -> Swift -> Nodejs -> Swift -> Firefox, even if they had conforming SON implementations (of the current specification)?
- seagreen 10y ago> not be accepted by a conforming SON parser (that did unicode normalization on incoming keys) I don't think RFC 7159 actually mentions Unicode normalization at all. If it was in there I think it would be mentioned in the section on string comparisons, but that just talks about going unescaped codepoint by unescaped codepoint: https://tools.ietf.org/html/rfc7159#section-8.3 https://tools.ietf.org/html/rfc7159#section-8.3 EDIT: Does Unicode normalization change over time? If so we definitely have to leave it out and just go codepoint by codepoint, because we don't want the validity of a Son document to change.