5 ms·
We're especially excited for this release because we discovered a decade old inconsistency with Go's notation of ASCII/UTF-8 that's now fixed in Go 1.19. A dem
by Zamicol 4y ago
We're especially excited for this release because we discovered a decade old inconsistency with Go's notation of ASCII/UTF-8 that's now fixed in Go 1.19.
A demonstration of the issue: https://go.dev/play/p/R9dm6uYZSas?v=goprev https://go.dev/play/p/R9dm6uYZSas?v=goprev
which fails in 1.18 but passes in 1.19.
A little background of the issue:
1. Go uses UTF-8 encoding. UTF-8 is a superset 7 bit, 128 code point ASCII. (Two UTF-8 creators, Rob Pike and Ken Thompson, are also two of Go's creators.)
2. All UTF-8 points over byte 128 must be encoded as two bytes. All code points 128 and below are encoded as a single byte.
3. Single byte code points between 129 and 256 are valid in Go, but not as UTF-8.
4. Go uses different notations to make the distinction between 129+ code points that are two bytes and 1-256 code points that are a single byte, notably the code points in the 129-256 range. Go represents single byte code points with the \x notation, "\xXX", and multi-byte code points with the \u notation, \uXXXX. All ASCII and single byte 129-256 codes should always be represented with "\x" notation.
The core of the issue is that Go was printing code point 128, \x7F, as \u007f. As previously said, the "\u" notation is reserved for valid UTF-8 points past the ASCII range, which are always encoded as a minimum of two bytes. \x7F is a valid single byte code point in the ASCII range and should not be printed in the "\u" notation.
We discussed this issue on an old thread, and Rob filed a Github issue https://github.com/golang/go/issues/52062 https://github.com/golang/go/issues/52062, and it was fixed within hours.
So the following line in the release note is us!
>Quote and related functions now quote the rune U+007F as \x7f, not \u007f, for consistency with other ASCII values.
How did we find this? We're fans of arbitrary base conversion (https://convert.zamicol.com https://convert.zamicol.com) and we discovered this issue while checking our work in our Go libraries. If encoded as two bytes, base conversion would be less useful.
Thanks Go team!
- morelisp 4y ago> Go was printing code point 128, \x7F That's code point 127, which has always been quite a special case as the only non-printable ASCII codepoint larger than the smallest printable codepoint. > All ASCII and single byte [128]-256 codes should always be represented with "\x" notation... the "\u" notation is reserved for valid UTF-8 points past the ASCII range... all valid UTF-8 "\u" notation characters must be a minimum of two bytes. Says who? I mean I guess some consistency (though, with what?) is nice, but you should be plenty prepared for the other format too. `strconv.Quote` is only meant to be equivalent to Go's source string literal representation and that allows both, and JSON only allows \u, and Python's source literals also allow both, and and and.... > UTF-8 points... single byte code points... multi-byte code points This is... at best, unnecessarily and fundamentally confusing language.
- Zamicol 4y ago>Says who? The playground example (https://go.dev/play/p/R9dm6uYZSas?v=goprev https://go.dev/play/p/R9dm6uYZSas?v=goprev) is a great example. All single bytes, 0-255, print as the printable character or if non-printable as \x except 127. There's nothing special about 127 to deserve this. Inversely, \u denotes multibyte for all code points except 127. Once again, why is 127 special? There's two possible fixes: note that 127 is special (even without a reason, but at least document it), or change the behavior to align with everything else. UTF-8 itself was a response in part to perceived arbitrary decisions made in other encodings; I'm not surprised that the second fix was preferred. Our chief concern was how many bytes were used in encoding, and that's when we ran into this issue. If not fixed, our tests in our library had to notate why 127 is special (because Go says so), or hope for a change. Now that it's fixed, there's no need for downstream documentation. It's a minor change, but now no one else ever has to spend the time we took to look into this issue because now there are no surprises. That makes it worth it. >That's code point 127 How does that joke go? There's only two hard things in computer science...
- morelisp 4y agoSo, says you, because it was aesthetically unpleasant for you. That's a far cry from "should always be... reserved for... must". Now we need to lockstep our team's version upgrade since I just learned some tooling will otherwise bounce it back and forth. "Thanks." > There's nothing special about 127 to deserve this. Other than the thing I said, which means it is now literally a special-case in the fmt code where it wasn't before.
- Zamicol 4y ago>Now we need to lockstep our team's version upgrade since I just learned some tooling will otherwise bounce it back and forth That sounds like an interesting issue. Could you perhaps go into more detail? >means it is now literally a special-case in the fmt code No. There is no special case, and the logic (literally) runs in the same switch case. https://go-review.googlesource.com/c/go/+/397255/4/src/strconv/quote.go#102 https://go-review.googlesource.com/c/go/+/397255/4/src/strco...