3 ms·
There's no definition of String.length that would be the obvious right choice. It depends on the use case. So probably better to let the application provide its
by winternewt 2y ago
There's no definition of String.length that would be the obvious right choice. It depends on the use case. So probably better to let the application provide its own implementation.
- josephg 2y ago> So probably better to let the application provide its own implementation. I’d be very happy with the standard library providing multiple “length” functions for strings. Generally I want three: - Length in bytes of the utf-8 encoded form. Eg useful for http’s content-length field. - Number of Unicode codepoints in the text. This is useful for cursor positions, CRDT work, and some other stuff. - Number of grapheme clusters in the text when displayed. These should all be reasonably easy to query. But they’re all different functions. They just so happen to return the same result on (most) ascii text. (I’m not sure how many grapheme clusters \0 or a bell is). Javascript’s string.length is particularly useless because it isn’t even any of the above methods. It returns the number of bytes needed to encode the string as UTF16, divided by 2. I’ve never wanted to know that. It’s a totally useless measure. Deceptively useless, because it’s right there and it works fine so long as your strings only ever contain ascii. Last I checked, C# and Java strings have the same bug.
- singpolyma3 2y agoDon't you want grapheme clusters for cursor positions? Length in encoded form can be found after encoding by checking the length of the binary content I guess. I think for historical reasons access to codepoints can be useful, but it's rarely what one wants.
- yau8edq12i 2y agoI don't know about Java, but the C# standard library is exceptionally well design with respect to variable byte encoding. https://learn.microsoft.com/en-us/dotnet/standard/base-types/character-encoding-introduction https://learn.microsoft.com/en-us/dotnet/standard/base-types... The built-in string.length method is useless (it returns the number of char objects) and I agree that's a problem, but the solution is also built into the language, unlike in JS.
- LegionMammal978 2y agoJS these days also has ways to iterate over codepoints and grapheme clusters. If you treat the string as an iterator, then its elements will be single-codepoint strings, on which you can call .codePointAt(0) to get the values. (The major JS engines can allegedly elide the allocations for this.) The codepoint count can be obtained most simply with [...string].length, or more efficiently by looping over the iterator manually. The Intl.Segmenter API [0] can similarly yield iterable objects with all the grapheme clusters of a string. Also, the TextEncoder [1] and TextDecoder [2] APIs can be used to convert strings to and from UTF-8 byte arrays. [0] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Intl/Segmenter https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... [1] https://developer.mozilla.org/en-US/docs/Web/API/TextEncoder https://developer.mozilla.org/en-US/docs/Web/API/TextEncoder [2] https://developer.mozilla.org/en-US/docs/Web/API/TextDecoder https://developer.mozilla.org/en-US/docs/Web/API/TextDecoder