2 ms·
On the opposite end of the spectrum from this idea, I've been thinking for a long time now that much of what's weird/redundant about Unicode comes from its self
by derefr 6d ago
On the opposite end of the spectrum from this idea, I've been thinking for a long time now that much of what's weird/redundant about Unicode comes from its self-synchronization requirement. And that you could create a very compact and flexible "re-embedding" of Unicode if you dropped this requirement.
Self-synchronization is the idea that if you take an arbitrary Unicode-encoding-encoded text stream, and then do any combination of 1. flipping bits in it at random, 2. injecting random extra whole bytes into the stream, and 3. dropping random whole bytes from the stream, then each such change will corrupt at most one Unicode code-unit in the stream, and from that, corrupt at most one semantic grapheme-cluster encoded by the stream. No such change will corrupt the stream itself so as to leave the stream in an invalid/indeterminate state that a Unicode parser can't know how to recover from. You'll have a one-character-wide "hole", and then the stream will resume. Unless that hole occurred at a character that's critical to the stream's meaning on an application-semantics level, the document will still be valid/useful (especially for archival/forensic-recovery purposes); just like a printed document is still valid/useful even if you drip a bit of ink on it.
Unicode encodings are self-synchronizing at the byte-pattern level. But, much less often discussed, Unicode itself is also designed to be self-synchronizing in how it encodes interactions between code-units. (In other words, Unicode is designed to never have modal or stateful semantics, beyond the boundary of a single grapheme cluster.) And this really constrains how certain Unicode features can be, and historically have been, designed.
If you want a run of Unicode code-units to all be "tagged" or "colored" with some property, then, due to the self-synchronization requirement (i.e. due to the assumption that any single one of those bytes could be corrupted or blown away, including whatever metadata-encoding bytes your scheme wants to use), you have to either:
- define a "pre-colored" alphabet, and express your tagged content in that alphabet. (Think of e.g. the Unicode flag codepoints [https://en.wikipedia.org/wiki/Regional_indicator_symbol https://en.wikipedia.org/wiki/Regional_indicator_symbol], which are essentially a special namespaced copy of the roman upper-case alphabet intended to be used only to spell out two-letter contiguous pairs that are [or at some point were] valid ISO country codes; where, when used in this way, the resulting 'colored'-letter-pair sequence has 'flag semantics', i.e. is meant to be rendered as a flag and machine-legible as a flag)
- or individually tag each and every one of those codepoints with its own tag/color codepoint (as in Unicode variation selectors)
- or interject binary-infix "operator" codepoints as glue between each codepoint in a codepoint sequence, so that those "operator" codepoints each affix together their immediate sibling codepoints, essentially constructing an abstract-Unicode-semantics list ADT "cons by cons", so that said list then may then be assigned its own semantics, e.g. being treated as a single grapheme-cluster with its own rendering (as in e.g. ZERO WIDTH JOINER used in its role in constructing complex emoji)
These encodings are all very high-overhead; but these are the kinds of trade-offs you have to make for self-synchronization to work.
And these trade-offs are sensible to make... if you're Ken Thompson in 1992, having to consider e.g. plaintexts being transmitted over raw RS232, or filesystems that just blast bytes to a spinning-rust disk without so much as a checksum, and then read them back "blindly" years later with that disk potentially highly-degraded.
But what if you live in the modern world, and you only care about holding and manipulating known-length strings in memory and/or embedded into code-signed (and thereby hashed) binaries; checksummed-block filesystems over rarely-corrupting NVMe; and transmission of data mostly over encrypted-stream protocols, where even the rare packet-level corruption that still TCP-checksums correctly, doesn't decrypt successfully, and therefore causes TLS-level retransmission?
Well, then you could define a much-more-concise reformulation of Unicode, that uses all the "forbidden" semantics-encoding techniques that Unicode itself avoids due to the self-synchronization constraint.
Where by "reformulation", I mean: a standard that keeps parity with Unicode in terms of what it can encode; and which at all times maintains a clearly-defined lossless bijective transformation between it and Unicode, evolving in lockstep with Unicode; but where Unicode and this formulation have their own distinct universes of codepoints, that compose using different rules, into the same ultimate sets of reachable grapheme-clusters with the same text-segmentation/collation/etc semantics.
I'm honestly kind of surprised that there isn't already a project somewhere to define an alternative Unicode formulation that looks like this.
It'd not only be potentially a highly-efficient representation for many in-memory string operations (that text shaping libraries would also love); it'd also likely perform impressively (compared to regular UTF-8) as a canonical representation for documents used to train+prompt LLMs. It'd give "more meaning per token", via all the repeated-per-codel overhead becoming once-per-sequence overhead; and it'd also allow many layers of meaning that are currently encoded via in-band protocols (ANSI escape codes, Markdown, HTML/XML, etc) that the LLM needs to learn additional recognition logic for, to instead be parsed out "during" initial text-stream recognition.
(And it's taking me real willpower not to go into depth on all the features such a formulation could have, and all the benefits it could provide. I should probably stop here before I nerd-snipe myself!)