5 ms·
the code unit size, byte order, and how many code units are in that codepoint. you can fit all of that information into 4 bits when you just care about minimiz
by bumblebritches5 8y ago
the code unit size, byte order, and how many code units are in that codepoint.
you can fit all of that information into 4 bits when you just care about minimizing the size of that mini header.
I know that, because I got bored and made up a quick proposal on it, not that I expect anyone else to even look at it.
- Dylan16807 8y agoIf you don't know code unit size and byte order up front, how do you even find the header? But I'm not really seeing the value. This is a variable-length format that isn't self-synchronizing, right? And if the header is fixed-length, then won't Latin letters require at least two bytes? It seems like it just loses to UTF-8 (which can support 31 bit code points, too).
- bumblebritches5 8y agobecause it's the top 4 bits... you read that as a byte, then read X bits more to get the rest of the code unit. yup and yup. UTF-8 is capped at 21 bits, theoretically it can support up to 6, but it's not allowed to because of UTF-16. This format would allow 28 bits to be encoded as 8, 16, and 32 bit code units. There would be no surrogate pair nonsense, simply write all the non-leading-zero bits into as many code units you need as well as the 4 bit header, and you're good. I'm not saying it'll ever happen, I know it won't, but if we could go back in time...
- Dylan16807 8y agoIf there's unknown byte order, then bytes might go 21436587 and you won't know if you're actually reading the byte that has the header. If we could go back in time, the important thing would be killing UCS-2, and the false idea of all code points fitting into 16 bits. The encodings with 8 bit and 32 bit code units are just fine, and don't encourage terrible assumptions. I'd definitely suggest a variable-width header, though. If you remove the extra bits that UTF-8 uses to be self-synchronizing, then your header is a fixed 1/8 overhead, which means you can fit 28-bit code points into four bytes while also fitting ASCII characters into one byte.
- bumblebritches5 8y agoThat's precisely what I meant with your UCS-2 comment. What do you mean by variable width? the 4 bit header would be limited to the first code unit, the following ones would just be pure codepoint values. and that's a good point, I as picturing it as the top 4 bits that describe the format, so the header's byte order would be field, but the actual codepoint value would be encoded with whatever byte and bit order the header indicated.
- Dylan16807 8y agoOh, so it's a BOM? Why integrate it into the code unit if you're only going to have one of them?
- bumblebritches5 8y ago> Oh, so it's a BOM? Kinda? It isn't limited to just the byte order tho, but I can see it. > Why integrate it into the code unit if you're only going to have one of them? There would be multiple code units, just 1 of the mini headers/BOM. With that comment I was trying to contrast it to UTF-8 which has those 2 leading bits of the continuation code units set to `0b10`, it wouldn't have that.
- Dylan16807 8y agoOh, the first code unit of each code point, not the first code unit of the text. That's how I had originally understood you until I got confused by the talk of byte order. I have to say I really don't see the value of being byte-order flexible, but always requiring the header be in the "first" byte, because that makes it so big-endian and little-endian formats are packed completely differently. 2 and 4 byte code units will be misaligned all the time too, might as well just use a series of bytes in fixed order. > What do you mean by variable width? Make the first byte be like UTF-8 (so a prefix of 0, 10, 110, or 1110), and the following bytes be raw data. It compacts a lot better than a fixed-length header, with roughly the same level of complexity. But now that I think about it a continuation bit is probably the best option here because it regains some ability to self-synchronize.