5 ms·
I actually read it as a argument FOR types and against modern languages choice to make the String class a weak proxy for typeless byte arrays. See all the argum
by arcbyte 6y ago
I actually read it as a argument FOR types and against modern languages choice to make the String class a weak proxy for typeless byte arrays. See all the arguments (in this HN comments no less!) for just using utf8 byte arrays as strings.
Hes saying semantically there's no difference between arrays and string classes except that with string classes we let you do all kinds of dangerous byte manipulation that we would never dream of with any other type. Moreover, most of the uses for this dangerous access aren't real usages because if you're manipulating strings you're almost certainly actually manipulating code points. So why wouldn't you just use a code point array and give yourself real type safety instead?
- ncmncm 6y agoI did not get that at all. Anyway a code point array would not serve the purpose: most possible sequences of valid code points are not valid strings. A variable-size array of code points is also useful, just as, in C++, a std::vector<char> is useful, but that doesn't make it a string. That C++ std::string<> is wrong for what we now think of as strings is a whole other argument. People once hoped that std::string<wchar_t> or std::string<char32_t> might be the useful string, but they were disappointed. C++ does not have a useful string type at this time, but there is ongoing work on one. It should appear in C++26.
- AnimalMuppet 6y ago> most possible sequences of valid code points are not valid strings. Could you clarify? In what way are they not valid strings?
- KMag 6y agoSome code points are characters. Others are operators with constrained contexts in which they operate. Sufficiently long random sequences of characters and these context-specific operators are likely to apply the operators in invalid contexts. Invalid characters mean invalid strings. For instance, there are code points that are effectively operators that add continental European accents (umlaut, accent grave, etc.) to Latin characters. (Also, there are redundant code points for accented characters.) There's a whole set of code points that are combinators for primitive components of Han characters, etc. (Also, there are redundant code points for pre-composed Han characters.) One way of writing Korean syllables strictly requires triplets of individual jamo components: initial consonant jamo, vowel jamo, and final consonant jamo. (Also, there are redundant code points for every valid triple-jamo syllable in Korean.) A Han character with an ancient Greek digamma in its "radical" position, a poo emoji inside a box, a thousand umlauts, all three French accents, a Hangul jamo vowel sticking through its center, a Hebrew vowel point, and a Thai tone mark is not a valid character. Any string containing invalid characters is not a valid string.