4 ms·
This discussion is about what encoding is best to choose for the low-level string types in languages and libraries, as that's what the original article was abou
by lambda 11y ago
This discussion is about what encoding is best to choose for the low-level string types in languages and libraries, as that's what the original article was about.
If your language already uses a Unicode-compatible string type, you should probably use that.
If you are designing a new language, rewriting a language's basic string primitives, or writing an API or application in a language that doesn't have a well-standardized Unicode string type like C or C++, then it is probably better to pick UTF-8; if your app is necessarily single platform and tied to a platform API, choose that API's encoding, but for anything cross platform, the encoding is going to differ between the platforms so choosing UTF-8 and then transcoding at the API boundaries on the different platforms is fine.
By the way, most languages do have the concept of a byte array, and you can provide UTF-8 string operations on top of that; but as I said, in general, you should be using the language's native string type unless you have a very good reason not to.
The original article that we're discussing does mention one of those very good reasons; in Python, it uses an 8-bit Latin-1 encoding if the characters can fit, upgrades to 16-bit UCS-2 if they can fit in that range, and finally upgrades to 32-bit UTF-32 if they don't fit in the UCS-2 range. But that means that if you're appending strings, and one of them contains an emoji (for instance, in a web templating system that encounters an emoji in a comment), that can mean that all of a sudden your string has to grow to four times its original size, which can cause serious problems; so one workaround for that may be to actually write everything out to a UTF-8 encoded string (which in Python can be represented as a byte array) rather than trying to append strings together and hitting that behavior.
Furthermore, several languages already do use UTF-8 as their internal string type; Rust, Go, and Ruby all use UTF-8 as their default or only string type.
And by the way, Rust and Go both use slices to handle the use case you describe; a slice is represented as a pointer to the first byte of interest, and a number representing how many bytes it extends. I don't know about in Go, but in Rust the API guarantees that all string slices created from a string begin and end on proper codepoint boundaries. Go: https://blog.golang.org/slices https://blog.golang.org/slices, Rust: https://doc.rust-lang.org/std/primitive.str.html https://doc.rust-lang.org/std/primitive.str.html.
- lisper 11y agoI think we're in violent agreement here. Keep in mind that the question I was responding to was: > Why would you want constant time access to code points? And the answer is: because you might want to compute an offset into a string in one place in your code, and then use that offset to access the string at that offset somewhere else in your code in constant time.
- lambda 11y agoBut you don't need constant time access to code points for that, you just need constant time access to code units, which UTF-8 can do just fine. That's why I had asked that question; every reason I have ever seen given for wanting constant time access by code point can either be done simply by indexing by code unit, or is misguided as it assumes that a single code point is an interesting unit of processing on its own. A code point (http://unicode.org/glossary/#code_point http://unicode.org/glossary/#code_point) is a value in the range from 0x0 to 0x10FFFF. In a fixed-width type, you need at least 21 bits to represent that, which is why 32 bits is the most common form for representing any arbitrary code point; but if you don't mind the alignment issues, you could even pack it to 24 bits or even 21 bits if you don't mind a little bit twiddling to extract it. A code unit (http://unicode.org/glossary/#code_unit http://unicode.org/glossary/#code_unit) is the smallest unit of encoding in the encoding that you are using. In UTF-8, that's a single byte; in UTF-16, it's two bytes, in UTF-32, four bytes. In order to be able to have constant time offsets, you only need to store a code unit index, not a code point index. The thing is, individual code points are not really all that interesting from the perspective of text processing. Almost all meaningful processing of them could potentially involve multiple code points, splitting on arbitrary code point boundaries without regards to higher level semantics can cause all kinds of problems, etc. So a desire to have constant time indexing by code points is misguided. All offsets can be stored by code unit (byte in UTF-8). That gives you all of the constant time access to given points of interest in a string that constant time indexing by code point would have given you.
- lisper 11y ago> But you don't need constant time access to code points for that, you just need constant time access to code units Well, sort of. What you really want is constant-time access to (semantically meaningful) substrings. Not all code points are semantically meaningful, but the situation is even worse for code units. A pointer to a random code point may or may not designate the boundary of a semantically meaningful substring. But whatever problems that causes, the situation for code units is even worse, because a pointer to a random code unit may or may not designate the start of a code point, and even if it does then you still have the previous problem that a random code point may or may not designate yada yada yada. (And yes, I know that UTF-8 is self-synchronizing. But that doesn't help with the problem I'm talking about, and if you think it does then you haven't understood the problem.) What you really want is constant-time dereferencing of designators for semantically meaningful substrings. But no language AFAIK actually has that. The fundamental problem is that most languages have painted themselves into a corner by carving into stone the fact that strings can be dereferenced by integers. Once you've done that, you're pretty much screwed. It's not that you can't make it work, it's just that it requires an awful lot of machinery. You basically need to build an index for every string you construct, and that can get very expensive. Fixed-width representations are a less-than-perfect-but-still-not-entirely-unreasonable engineering solution to this problem.