3 ms·
But UTF-8 is just a way to encode a number as a variable-length string of octets. Why would you be unable to encode, say, a terminating U+D800 as a string of th
by deadbeeves 3y ago
But UTF-8 is just a way to encode a number as a variable-length string of octets. Why would you be unable to encode, say, a terminating U+D800 as a string of three bytes at the end of a UTF-8 stream?
- skitter 3y agoBecause that's how UTF-8 is defined[1]. WTF-8 lifts that restriction. [1] https://simonsapin.github.io/wtf-8/#utf-8 https://simonsapin.github.io/wtf-8/#utf-8
- deadbeeves 3y agoIt doesn't sound very annoying, then. You use the exact same encoding scheme, but skip a verification step. Actually it sounds more convenient.
- jraph 3y agoStill potentially annoying if you deal with some other code that expects UTF-8 proper and you pass it a wtf-8 string that fails the lifted verification.
- deleted 3y ago[deleted]