4 ms·
But you don't need constant time access to code points for that, you just need constant time access to code units, which UTF-8 can do just fine. That's why I ha
by lambda 11y ago
But you don't need constant time access to code points for that, you just need constant time access to code units, which UTF-8 can do just fine. That's why I had asked that question; every reason I have ever seen given for wanting constant time access by code point can either be done simply by indexing by code unit, or is misguided as it assumes that a single code point is an interesting unit of processing on its own.
A code point (http://unicode.org/glossary/#code_point http://unicode.org/glossary/#code_point) is a value in the range from 0x0 to 0x10FFFF. In a fixed-width type, you need at least 21 bits to represent that, which is why 32 bits is the most common form for representing any arbitrary code point; but if you don't mind the alignment issues, you could even pack it to 24 bits or even 21 bits if you don't mind a little bit twiddling to extract it.
A code unit (http://unicode.org/glossary/#code_unit http://unicode.org/glossary/#code_unit) is the smallest unit of encoding in the encoding that you are using. In UTF-8, that's a single byte; in UTF-16, it's two bytes, in UTF-32, four bytes.
In order to be able to have constant time offsets, you only need to store a code unit index, not a code point index.
The thing is, individual code points are not really all that interesting from the perspective of text processing. Almost all meaningful processing of them could potentially involve multiple code points, splitting on arbitrary code point boundaries without regards to higher level semantics can cause all kinds of problems, etc. So a desire to have constant time indexing by code points is misguided.
All offsets can be stored by code unit (byte in UTF-8). That gives you all of the constant time access to given points of interest in a string that constant time indexing by code point would have given you.
- lisper 11y ago> But you don't need constant time access to code points for that, you just need constant time access to code units Well, sort of. What you really want is constant-time access to (semantically meaningful) substrings. Not all code points are semantically meaningful, but the situation is even worse for code units. A pointer to a random code point may or may not designate the boundary of a semantically meaningful substring. But whatever problems that causes, the situation for code units is even worse, because a pointer to a random code unit may or may not designate the start of a code point, and even if it does then you still have the previous problem that a random code point may or may not designate yada yada yada. (And yes, I know that UTF-8 is self-synchronizing. But that doesn't help with the problem I'm talking about, and if you think it does then you haven't understood the problem.) What you really want is constant-time dereferencing of designators for semantically meaningful substrings. But no language AFAIK actually has that. The fundamental problem is that most languages have painted themselves into a corner by carving into stone the fact that strings can be dereferenced by integers. Once you've done that, you're pretty much screwed. It's not that you can't make it work, it's just that it requires an awful lot of machinery. You basically need to build an index for every string you construct, and that can get very expensive. Fixed-width representations are a less-than-perfect-but-still-not-entirely-unreasonable engineering solution to this problem.
- lisper 11y agoI wrote: > (And yes, I know that UTF-8 is self-synchronizing. But that doesn't help with the problem I'm talking about, and if you think it does then you haven't understood the problem.) I would like to retract this, and apologize for the snarky tone. I had mis-remembered the design of UTF-8, and so pretty much everything I said was wrong (well, not all of it was wrong, but the parts that weren't were exactly what you were saying). I'd edit or delete my comment if I could, but the deadline has passed so this is the best I can do.