3 ms·
From the article: >UTF-16 is designed to represent any Unicode text, but it can not represent a surrogate code point pair since the corresponding surrogate 16-
by j_jochem 11y ago
From the article:
>UTF-16 is designed to represent any Unicode text, but it can not represent a surrogate code point pair since the corresponding surrogate 16-bit code unit pairs would instead represent a supplementary code point. Therefore, the concept of Unicode scalar value was introduced and Unicode text was restricted to not contain any surrogate code point. (This was presumably deemed simpler that only restricting pairs.)
This is all gibberish to me. Can someone explain this in laymans terms?
- cygx 11y agoPeople used to think 16 bits would be enough for anyone. It wasn't, so UTF-16 was designed as a variable-length, backwards-compatible replacement for UCS-2. Characters outside the Basic Multilingual Plane (BMP) are encoded as a pair of 16-bit code units. The numeric value of these code units denote codepoints that lie themselves within the BMP. While these values can be represented in UTF-8 and UTF-32, they cannot be represented in UTF-16. Because we want our encoding schemes to be equivalent, the Unicode code space contains a hole where these so-called surrogates lie. Because not everyone gets Unicode right, real-world data may contain unpaired surrogates, and WTF-8 is an extension of UTF-8 that handles such data gracefully.
- SimonSapin 11y agoEvery term is linked to its definition. https://simonsapin.github.io/wtf-8/#terminology https://simonsapin.github.io/wtf-8/#terminology Does this help?
- haberman 11y agoThis was gibberish to me too. I researched it a bit and wrote an explanation that would have made sense to the 2-hours-ago me: https://news.ycombinator.com/item?id=9614641 https://news.ycombinator.com/item?id=9614641