4 ms·
Treating Unicode strings as a sequence of code points is a completely valid thing to do, but is usually not what you actually care about when dealing with text.
by pjscott 3y ago
Treating Unicode strings as a sequence of code points is a completely valid thing to do, but is usually not what you actually care about when dealing with text. Really, are code points any less of an implementation detail?
- croes 3y agoI think parent means most correct of the three given examples from Rust, JS and Python. Especially because the article says that Python's take is the worst.
- paulddraper 3y agoYes. They are less of an implementation detail. Grapheme > Code point > Encoding > Endianness > Media It's all "implementations" but some are lower then others
- planede 3y agoCode points are what you care about when you do any kind of text-based format encoding or decoding. Any of JSON, XML, HTML, YAML or whatever is defined by sequence of code points. There is no reason to complicate these with visual representation-specific concepts. If you have to care about the visual representation of text then you probably need to be familiar with other concepts as well.
- chrismorgan 3y agoBut, given the root ancestor of this comment, it’s worth clarifying that Python’s approach to strings doesn’t help at all with things like decoding JSON/XML/HTML/YAML; what Python gives you is random access by code point index, which you won’t ever need to use in such tasks.