3 ms·
Self-synchronization in UTF-8 is intuitively a great thing to have, yet I don’t remember actively relying on it ever. Does anyone have a good example of when it
by lukasgelbmann 13d ago
Self-synchronization in UTF-8 is intuitively a great thing to have, yet I don’t remember actively relying on it ever. Does anyone have a good example of when it‘s useful?
Another related nice property that UTF-8 has: substring search reduces to bytestring substring search. I.e. given two Unicode strings in UTF-8 encoding, you can check if one is a substring of the other by just treating them as bytestrings and checking if one bytestring is a substring of the other bytestring. This is a stronger property than self-synchronization: UTF-8 has it, but UTF-8000 doesn’t.
- hippietrail 13d agoOnly if they've already both undergone normalization to NFC or NFD.
- conradludgate 13d agoRely on it? Not that I can remember. However, Rust makes use of it for fast safety checks. Because rust strings must be valid utf8, if you want to take a substring at some range, eg "Hello, World!"[7..12] then it's very simple to just check bytes 7 and 12 and see if they are the start of a codepoint, no other scanning or parsing is required.
- tjjfvi 13d agoI don’t see how UTF-8000 doesn’t have it. The first byte of any code point is either 0xxxxxxx or 10xxxxxxx, which is distinct from all non-first bytes which are 11xxxxxx. Thus any UTF-8000 sub-bytestring must necessarily have the start aligned at a code point boundary, at which point all the subsequent bytes are interpreted as codepoints in the same way.
- lukasgelbmann 13d agoRight, UTF-8000 does have this property. Too late to edit my comment now, thanks for noticing that.
- deleted 13d ago[deleted]
- layer8 13d ago> Does anyone have a good example of when it‘s useful? It prevents vulnerabilities where an incorrect offset into a string could result in characters being read that aren’t in the original string (which could defeat a prior sanitization of the string).
- rurban 13d agoOnly if its normalized unicode. Most strings are not, and utf-8 does not guarantee normalization.
- yencabulator 13d agoI can think of 3 things: 1. You can partition an input file at any offsets, parallelize, and adjust partition boundaries to a valid offset independently. Without the property, parallelization is hard. This is how mapreduce has been used to process large text files, except at line boundaries. Now, for this that might not be a useful enough property, given that we already do similar things for newlines, and UTF-8 guarantees ASCII is always recognizable and hence newlines are always recognizable. 2. It might have been more useful in the era of dial-up where we still had occasional corrupted bytes in the transmission. 3. It helps regain sanity if e.g. a background process outputs bytes that get interleaved at the tty. For example, cat a large text file, the write boundaries won't always align at UTF-8 boundaries, then have a background process output get interleaved in an unfortunate way. If it self-synchronizes, it'll knock itself back into sync after a small amount of garbage.