3 ms·
JSON[1], for instance, specifies: 1. "JSON syntax describes a sequence of Unicode code points. JSON also depends on Unicode in the hex numbers used in the \u e
by _rend 5y ago
JSON[1], for instance, specifies:
1. "JSON syntax describes a sequence of Unicode code points. JSON also depends on Unicode in the hex numbers used in the \u escapement notation."
2. "A JSON text is a sequence of tokens formed from Unicode code points that conforms to the JSON value grammar."
3. "A string is a sequence of Unicode code points wrapped with quotation marks (U+0022). All code points may be placed within the quotation marks except for the code points that must be escaped: quotation mark (U+0022), reverse solidus (U+005C), and the control characters U+0000 to U+001F. There are two-character escape sequence representations of some characters."
4. "Any code point may be represented as a hexadecimal escape sequence. The meaning of such a hexadecimal number is determined by ISO/IEC 10646. If the code point is in the Basic Multilingual Plane (U+0000 through U+FFFF), then it may be represented as a six-character sequence: a reverse solidus, followed by the lowercase letter u, followed by four hexadecimal digits that encode the code point."
5. "Note that the JSON grammar permits code points for which Unicode does not currently provide character assignments."
JSON does require Unicode awareness, both for general parsing, and for correctly interpreting strings. Backslashes are allowed for special escape characters, which means that you must be aware of the format (and the encoding of the text) in order to be able to decode.
Note that JSON also doesn't specify a required encoding, only Unicode correctness, so a parser may need to be able to handle multiple Unicode encodings and differentiate between them.
The spec doesn't specify what to do with code points which are not understood as Unicode (given especially the allowance for unassigned characters), but explicitly-invalid Unicode should be rejected.
[1] https://www.ecma-international.org/wp-content/uploads/ECMA-404_2nd_edition_december_2017.pdf https://www.ecma-international.org/wp-content/uploads/ECMA-4...
- Spex_guy 5y agoParsing this out of utf-8 encoding requires no knowledge of unicode or even utf-8. All of the relevant characters (reverse solidus, quotation mark, and control characters) are single byte characters in the ascii subset. These characters cannot be found inside multi-byte characters in utf-8 due to the design of the encoding. Converting the unicode character escape codes to utf-8 would require knowledge of utf-8 encoding, but this unescaping is not a feature that would be provided by the language regardless.
- _rend 5y ago> Parsing this out of utf-8 encoding requires no knowledge of unicode or even utf-8. If you have valid UTF-8 already, then yes, the task is a lot easier. But depending on the level at which you're parsing, this might not be the case — i.e., if you're writing a JSON parser from the ground up, you do need to know what UTF-8 and Unicode are, and will need to validate the input data. > Converting the unicode character escape codes to utf-8 would require knowledge of utf-8 encoding Agreed. Even if you're not working at the "array-of-bytes" level, you will need to be able to parse and translate "\u..."-style strings into the appropriate output character encoding. > but this unescaping is not a feature that would be provided by the language regardless. I'm not sure we're talking about this being handled at the language level. This translation is something that would likely be offered at the parser level (working with the features offered by the standard library), but the parser does need to know about it — and does need to be able to work with strings at a granular level to be able to parse it out. By definition, it cannot leave the input data as an undecoded bag of bytes. Note, too, that the JSON spec does not specifically require UTF-8. UTF-16 is a completely valid encoding for JSON (though much less common than UTF-8), in which case none of these characters are an ASCII subset, and greater awareness is needed to be able to handle this.
- lhorie 5y ago> it cannot leave the input data as an undecoded bag of bytes But all it's doing here is taking a hex string (which is entirely ASCII) and converting it into the respective hex representation. Since ASCII translates unambiguously to bytes, it doesn't really matter if `str[0]` is operating on a byte stream, codepoint stream or grapheme stream, because in utf8, they're all the same thing as long as we're within the ASCII range. Where things get hairy is stuff like `str.reverse()` over arbitrary strings that may or may not be in ASCII. This repo[0] talks about some of the challenges associated with conflating characters with either bytes or codepoints. The problem is that programming languages often approach strings from the wrong angle: you can't just tack on handling of multi-byte codepoints on top of ascii handling; you lose O(1) random access and you don't actually model the linguistic domain properly by doing so, because in the first place, humans think of characters not in terms of bytes or codepoints, but in terms of grapheme clusters. Clustering correctness falls deep in the realm of linguistics, and is therefore arguably more suitable to be handled by a library than a programming language. [0] https://github.com/mathiasbynens/esrever https://github.com/mathiasbynens/esrever