4 ms·
The biggest takeaway/shock for me in this article is the fact that JSON string literals can't contain escaped characters outside the Basic Multilingual Plane (i
by poorlyknit 3y ago
The biggest takeaway/shock for me in this article is the fact that JSON string literals can't contain escaped characters outside the Basic Multilingual Plane (i.e. whose code points are greater than U+FFFF).
Quoting RFC 8259:
To escape an extended character that is not in the Basic Multilingual
Plane, the character is represented as a 12-character sequence,
encoding the UTF-16 surrogate pair.
So in order to encode U+1F914 THINKING FACE you can either do
{"text": "<thinking face>"}
(Emoji omitted bc of HN) or
{"text": "\ud83e\udd14"}
but not
{"text": "\u1f914"}.
This seems to be a relic from ECMAScript which (iirc, only skimmed it) assumes UCS-2/UTF-16. From that perspective it makes a lot of sense but it steel feels a little icky to me having surrogates referenced in a standard that is supposed to be UTF-8 only :)
- chubot 3y agoRight exactly. The conventional ways of writing it would be \U0001F914 -- must be exactly 8 digits \u{1f914} -- 1 to 6 digits Almost all languages have these now, but JavaScript didn't when JSON was invented. The "J8 strings" extension mentioned technically only needs \a syntax for bytes, which is \yff since \xff is (oddly) a synonym for \u00ff and thus unsuitable. But I also want to add \u{123456} because it allows people to move away from the weird UTF-16 legacy in a UTF-8 format.
- poorlyknit 3y agoYea, it seems weird to me that they decided to do that in ECMAScript in the first place. The user/programmer should never have to interface with surrogates unless they are doing something specific to UTF-16.
- masklinn 3y agoJavascript was released about 6 months before unicode 2.0, and furthermore was in various ways designed to be close to Java, whose strings are sequences of 16 bits code units. Because this is part of the string interface it’s not really fixable. Not only that but while Unicode 2.0 was released in June 1996, it wasn’t necessarily super sought after for a while, lots of people were wary of utf8 and while 16 bit code units was fine 32 was a bit much. While it was singularly late to the party, it took MySQL until 2010 to support non-BMP content, unless you were willing to store content as blobs. Java and Javascript similarly took a while just to give access to actual codepoints, in Java5 (2004, String#codePointAt, it took until Java 8 for an iterator to be added with CharSequence#codePoints) and ES6 (2015, String.prototype.codePointAt and String.prototype[@@iterator])
- TRiG_Ireland 3y agoAs JSON must be in Unicode (I think the latest RFC restricts it to UTF-8), there's no actual _need_ to escape emoji. But if you do need to (perhaps you're in an ASCII-only environment), you have to do it that way, yes. Weird. (I asked once, on Stack Overflow, why this was the case, and was merely told that it was historical reasons, which matches what you surmised: it's a relic.)
- poorlyknit 3y agoRight! In practice you'd just put the symbols directly. Funnily enough {"text": "\ud83e\udd14"} is the only way to encode the JSON from my post on HN :D
- dataflow 3y agoWouldn't \u1f914 have been interpreted as \u1f19 followed by '4'? How were you imagining this might work?
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]