4 ms·
Can you give an example of how one might make use of the last code point in a string?
by svrb 6y ago
Can you give an example of how one might make use of the last code point in a string?
- rectang 6y agoI tried to simplify the example code for the sake of clarity. While I acknowledge it would be unusual to provide a library API for finding the value of the last code point specifically, most string libraries provide something like a "codePointAt" function to return the code point at a specific offset. You'll run into the same problem with DECODE_UTF8 there. int32_t code_point_at(const char *s, size_t len, size_t offset) { const char *end = s + len; while (s < end && offset > 0) { // Move forward by 1-4 bytes depending on header byte value. s += UTF8_SKIP(s); offset--; } if (s >= end) { return -1; } return DECODE_UTF8(s); // BOOM! May read beyond string end }
- oconnor663 6y agoIn practice, this sort of thing might be most likely to come up with a chars iterator. The natural iteration logic is something like: 1) If position == end, the iterator is done. 2) Otherwise, read the next byte to determine the number of bytes in the next character. 3) Read all the bytes of the next character, convert them to a regular 32-bit code point, and emit that. In the presence of invalid UTF-8, that iterator is broken. There needs to be an extra check in step 2.5: "Check that the number of bytes reported for the next character doesn't exceed the number of bytes remaining in the string." Iterating over the characters of a string is a pretty common thing to do, and that extra branch is not free.
- qppo 6y agoThe check only needs to be done once, when the iterator is constructed. You can also zero pad the buffer by 3 bytes to avoid a bounds check if you really want. This seems like a premature optimization, unless you're only iterating over tiny strings.
- oconnor663 6y agoThat's a good point. I know in the case of Rust in particular there are other invariants, like that each 32-bit code point is guaranteed to have a valid value. (The compiler knows this and may use unset bits to stash enum state or something like that.) But maybe in other languages without such strict validity guarantees, it's less of a big deal?
- jart 6y agoIf the natural logic is broken then you've misunderstood the thing's nature. UTF-8 behaves more like a communications stream that was shoehorned into the purpose of character arrays. If you think about it in that way, then decoding can be done simply and elegantly: #define bsr(u) ((sizeof(int) * 8 - 1) ^ __builtin_clz(u)) #define ThomPikeByte(x) ((x) & (((1 << ThomPikeMsb(x)) - 1) | 3)) #define ThomPikeMsb(x) (((x)&0xff) < 252 ? bsr(~(x)&0xff) : 1) void ThompsonPikeDecoder(const char *s) { unsigned c, w = 0, t = 0; do { c = *s++ & 0xff; if (0200 <= c && c < 0300) { w = w << 6 | c & 077; } else { if (t) { printf("%04x\n", w); t = 0; } if (c < 0200) { printf("%04x\n", c); } else { w = ThomPikeByte(c); t = 1; } } } while (c); } That code generalizes to any 32-bit number. It's the full superset of intended behaviors including things like stream synchronization. It'll even decode numbers that were arbitrarily banned by the IETF e.g. \300\200. Most importantly, since it doesn't require a validation pass beforehand, does that mean it goes faster than the OP's code in praxis? D:
- Dylan16807 6y agoI strongly question "doesn't require a validation pass". 1. Your code drops naked continuation bytes entirely. 2. If a valid multibyte character is followed by extraneous continuation bytes it will not decode correctly. 3. Using the wrong number of continuation bytes is a giant mess and can give you arbitrary numbers in multiple ways. The second one especially is a behavior that's hard to defend.