3 ms·
In practice, use cases don't justify strings precalculating their number of code points or grapheme clusters. Pixel width isn't a property of a string alone, si
by hsivonen 9y ago
In practice, use cases don't justify strings precalculating their number of code points or grapheme clusters. Pixel width isn't a property of a string alone, since it depends on font and the container to render it in (for line breaks).
So a string should know its number of code units. Validity information should be part of the type system and not runtime flags. I.e. you have a different type for arbitrary bytes and for valid UTF-8. (See IDNA for why normalization guarantee in infra over time is problematic.)
- mjevans 9y agoHaving thought about it you're probably right. That properly belongs on a different kind of wrapper around this data structure or ones like it. Maybe in that /specific/ and /rare/ use case it even makes sense to store the underlying data in a wide data structure; but my supposition is that handling it as a series of fragments (which would have a length in memory and in code-points sizes) would be the answer. If addressing in such a way replacement and other editing operations seem likely.