3 ms·
This is actually mentioned in the submission, but it does so through surrogate pairs. In short: the first of two bytes references a different "plane" of code po
by devadvance 7y ago
This is actually mentioned in the submission, but it does so through surrogate pairs. In short: the first of two bytes references a different "plane" of code points, and the second byte refers to a code point in that additional plane.
What's cool is that UCS-2 and UTF-16 behave the same way in this regard, so you can see how this works in JavaScript as well. There's a great HN post about the length of an emoji being 2 in JS [1] from 2017.
Practically speaking, I had fun learning this while writing an Angular app to read SMS backups [2]. It's frustrating and cool at the same time.
[1] https://news.ycombinator.com/item?id=13830177 https://news.ycombinator.com/item?id=13830177
[2] https://github.com/devadvance/sms-backup-reader-2/blob/master/README.md https://github.com/devadvance/sms-backup-reader-2/blob/maste...
- deleted 7y ago[deleted]
- Thorrez 7y agoIf there are surrogate pairs, I thought that was more of UTF-16 rather than UCS-2. Wikipedia says UCS-2 cannot go outside the BMP range (which emojis are). >UCS-2 differs from UTF-16 by being a constant length encoding[4] and only capable of encoding characters of BMP. https://en.wikipedia.org/wiki/UTF-16 https://en.wikipedia.org/wiki/UTF-16 There does seem to be some disagreement on this topic though. I wonder if it could be described as WTF-16. https://simonsapin.github.io/wtf-8/ https://simonsapin.github.io/wtf-8/
- temac 7y agoThe submission actually states (correctly): "UCS-2 can only represent the first 65536 characters of Unicode, also known as the Basic Multilingual Plane." UTF-16 can represent all Unicode code points, through surrogate pairs when needed. Systems that originally were designed for UCS-2 have often been upgraded to UTF-16. E.g. Windows, JS. But it is then not UCS-2 anymore. In some case the documentation and/or function names, etc., might talk incorrectly about being UCS-2 (while they really are UTF-16 now) for historical reasons, adding to the confusion.