4 ms·
I think you may be confused about the meaning of "plane" in Unicode. A plane is a contiguous set of 65,536 (216) code points. The lowest-numbered plane was ori
by ubernostrum 8y ago
I think you may be confused about the meaning of "plane" in Unicode.
A plane is a contiguous set of 65,536 (216) code points. The lowest-numbered plane was originally the only one; now it's simply Plane 0 (officially, the "Basic Multilingual Plane"). "Astral Plane" is a collective term to refer to planes beyond that, of which there are now 16, for a total of 17 planes.
Unicode currently promises to cap at 17 planes because that's the limit of what UTF-16 can encode with surrogate pairs; anything beyond Plane 16 (officially, "Supplementary Private Use Area-B") would require some other scheme to encode in 16-bit units.
Unicode does not use anywhere near the amount of space even those 17 planes represent; you could encode seven full copies of current Unicode in the unassigned and non-reserved space available, and still have room left over (eight full copies if you un-reserve Planes 15 and 16).
If we ever truly needed more than 17 planes, we'd have trouble trying to do UTF-16, but that's a problem with UTF-16, not with Unicode. UTF-8 as currently defined would run into trouble beyond 32 planes, but we'd either abandon it for something else, or bolt on some inelegant hack to extend UTF-8 beyond four bytes.
So other than maintaining compatibility with existing encodings, there is nothing about Unicode's design that forbids tacking on more planes if and as needed. The hard part, as with so many things in programming, was going from "there's only one" to "there's more than one". If we ever needed enough planes to exhaust both UTF-16 and UTF-8.
One very obvious jest is that at the current rate of emoji expansion it may be inevitable to open up the next plane just for emoji.
To put it in perspective, the total set of emoji in Unicode 11.0, which are spread out across multiple blocks in different planes due to historical reasons, add up to less than 2% of the capacity of a single plane, and 0.11% of the available space in a 17-plane Unicode.
- filmor 8y agoThe scheme used in UTF8 can encode up to 42 bits with a single start byte, the last start byte being 0xFF followed by 7 bytes of the form 10xxxxxx.
- ubernostrum 8y agoHence I said "UTF-8 as currently defined". You could produce something that uses the same basic scheme as UTF-8 (using the leading byte to indicate the total number of bytes used for the code point), but it would not be UTF-8 as we know it (which caps at four bytes per code point), and different encoders/decoders would need to be developed.
- WorldMaker 8y agoI appreciate the technical description of a plane. I was indeed using the more colloquial "plane [collection]" as shorthand for "collection of planes" as in the "Astral Plane [collection]".