4 ms·
Why would you want constant time access to code points? What can you do with constant time access to code points that is actually correct? Every time someone t
by lambda 11y ago
Why would you want constant time access to code points? What can you do with constant time access to code points that is actually correct?
Every time someone tries to promote UTF-16, UTF-32, or some imagined encoding like UTF-24, they bring up constant time access to code points, but I have never heard of a reason why you would want that.
Pretty much every use case I have ever heard for constant-time access can be handled just as well by constant-time access at the code unit (byte in UTF-8, 16 bit value in UTF-16, 32 bit value in UTF-32) level. Matching text exactly? Matching works just as well at the code unit level. Doing any kind of fuzzy or regular expression match? You're going to need to iterate over each item anyhow to normalize it or classify it. Need to store offsets? That works just fine at the code unit level.
Most of the other suggestions I've heard for what to do with constant-time access to code points is incorrect when you consider combining characters, normalization, different character widths, need for linguistically appropriate word splitting, etc. If you're doing something that doesn't take these into account, why are you using Unicode instead of just ASCII?
In addition, for the ASCII range, UTF-24 would take up 3 times the space as UTF-8, and the vast majority of text processed is actually in the ASCII range due to verbose markup formats like HTML, XML, etc. Plus if you do anything with the code points in UTF-24, you need to do a bunch of bit-fiddling to move them into 32 bit alignment, so as far as actually decoding the individual code points, it's pretty much a wash with UTF-8.
- lisper 11y ago> What can you do with constant time access to code points that is actually correct? Compute an offset into a string in one part of your code, and then pass a representation of that offset to another part of your code that then accesses that part of the string and does something with it. It is theoretically possible to produce an abstract representation of such an offset that would allow constant-time access into a non-uniformly-represented string, but no programming language that I know of has this feature. Even Common Lisp, which is probably the most flexible and future-proof language ever designed, requires that offsets into strings be integers. (Actually, it's even worse than that: Common Lisp requires that strings be arrays of characters, which is truly ironic if you think about it.)
- rectang 11y agoBut as lambda points out, that offset could be measured in code units rather than code points. You still don't need random access into variable-width data.
- lisper 11y agoRead what I wrote more carefully. You have to take that offset, represent it somehow, and then pass that representation off to another part of the code, which then uses that representation to access the part of the string that was found earlier. How do you do that without forcing a re-scan when the only facility that the language provides for accessing a character in a string requires you to represent the location of that character as an integer? Actually, you can do it in C because you can use a raw pointer, but that is fraught with all manner of other perils.
- jcranmer 11y agoJust use the octet index instead of the code point index. To use UTF-8-based strings, you do need to distinguish between octet indexes (calling them iterators is probably saner to most people) and character indexes. With UCS-4 (and UCS-2 as implemented in practice), the iterator representation and character index are the same, which makes retrofitting those implementations to use UTF-8-based strings impossible. But for new environments, where UCS-2-backwards-compatibility isn't a requirement, it's not hard to make a string API that properly exposes the difference between iterators and character indexes, which is why newer environments are gravitating towards UTF-8 strings.
- maxlybbert 11y ago> Why would you want constant time access to code points? What can you do with constant time access to code points that is actually correct? Well, it makes it easier to avoid putting an ellipsis in the middle of a multi-byte character, which could crash the user's phone when something else tries to display your now-invalid string.
- lambda 11y agoNote that my question included "...that is actually correct" as a condition. Truncating between code points to insert an ellipsis is not correct. Not all characters have the same width, even in a fixed-width font; for instance, combining characters have zero width. Truncating between a base character and a following combining character makes no sense. If you're truncating the word "exposé" to fit into six characters, it makes no sense to count that as 7 characters, truncate the accent, and wind up with "expose...". (Yes, normally that will be represented as a pre-composed character, but there are base-character/diacritic combinations that have no pre-composed character, so you will necessarily get a combining diacritic). Truncating based on character count can also mean that you exceed your width requirement; CJK characters are twice the width of most alphabetic characters. Additionally, most text is displayed using variable width fonts, where truncation should depend on the precise sum of the widths of characters, but even text that isn't can have the previous two issues. These mean that to do truncation properly, you need to iterate over the characters counting width until you exceed your required width in order to truncate, then truncate at the previous grapheme boundary, which requires walking through the string in order and can be done just as efficiently in a variable width encoding as in fixed.
- maxlybbert 11y ago> Note that my question included "...that is actually correct" as a condition. I think it's valid to pick an approach that makes it more likely I'll do things correctly.
- vorg 11y ago> UTF-24 would take up 3 times the space as UTF-8, and the vast majority of text processed is actually in the ASCII range due to verbose markup formats like HTML, XML, etc Sounds like these "verbose" markup formats are the real problem in taking up too much space, not the Unicode transformation formats.
- Jweb_Guru 11y agoOkay. But that's the world in which we live. Complaining about the real problem being "things people put in strings" rather than the actual encoding format isn't going to solve any problems people actually have.
- lambda 11y agoI think that moving the world away from HTML and XML as markup languages is likely to be considerably harder than simply defaulting to using UTF-8 in any new code for which there isn't already a natural encoding to choose based on the platform or APIs you're developing on. Furthermore, even beyond the markup, the vast majority of written content is in the Roman alphabet, of which most characters are in the ASCII range, and those that aren't fit into 2 bytes of UTF-8. 55% of text on the web is in English, then languages using the Roman script or other scripts in the 2-byte UTF-8 range make up 21 of the next 25 most commonly used languages on the internet; of the top 25 languages, only Chinese, Japanese, Korean, and Thai are in ranges that require 3 bytes in UTF-8 for the majority of their characters (https://en.wikipedia.org/wiki/Languages_used_on_the_Internet https://en.wikipedia.org/wiki/Languages_used_on_the_Internet). So for the vast majority of text that is processed, UTF-8 is dramatically more efficient than a hypothetical UTF-24 or the real UTF-32. Now, you might say "well, it should all compress away anyhow", but when dealing with text in RAM, it is generally not compressed (and that would make handling it far more complex than just using UTF-8), and memory capacity, latency, and bandwidth can all be important. Trying to deal with text as fixed-width characters is simply incorrect. Almost every non-trivial means of handling text needs to support arbitrary length strings, needs to deal with clusters of more than one codepoint as a single unit, and needs to iterate over the characters linearly at least once (and can then store byte offsets for random access later on). Dramatically increasing storage requirements of text for the non-benefit of being able to deal with fixed-width codepoints just doesn't make sense.