14 ms·
We don't need a string type (2013)
- pca006132 6y agoI think the problem is that, a lot of time when we deal with strings, we are thinking about ASCII strings instead of other encoding like UTF-8. If we treat them as ASCII strings, an array of characters would make sense, but it is not that simple for other encoding. One of the languages that considered the issue is Rust. In rust, we don't really index into strings, but use iterators or other methods to do the operations required. https://doc.rust-lang.org/std/string/struct.String.html https://doc.rust-lang.org/std/string/struct.String.html
- sfvisser 6y agoI really don’t think many programmers nowadays actually think this.
- bobthepanda 6y agoI would hazard that very few people think about what an underlying String is at all. String encoding is something I encountered as a problem in college, but is up there with implementing a homemade red-black tree in terms of “things that are asked in interviews but have little to no bearing on my day-to-day.”
- BlueTemplar 6y agoReally, they don't run into string/character issues regularly ? Because I do...
- bobthepanda 6y agoI certainly run into them rarely, and if I do have an issue it is usually solved by bunging it into some purpose built standard or third party library and calling it a day. I’m sure people have jobs that deal with this, but the low-level form of the problem is not something that I could see one encountering in a meaningful way for building a standard CRUD app or service.
- belval 6y agoI completely agree with you, but everyone who doesn't have to deal with unicode/strings on a regular basis should consider themselves lucky. Once your add RTL text (with the matching bidi algorithm) or grapheme-based written system such as Devanagari which doesn't really have characters at all it becomes such a mess so fast.
- DougBTX 6y agoThe date should be (2013) not (2018), as that dates it before Rust 1.0 (which does have a UTF-8 string type) and before the Julia 1.0 release date (which implements UTF-8 strings as arrays with irregularly spaced indexes, eg, the valid indexes may be 1, 2, 4, 5, if the character at 2 takes up two bytes). Both would be interesting examples to compare against if this article was written today.
- dang 6y agoI've fixed the date now. Actually the date at the top of the article "2013-08-13" is in a font that somehow makes it look like 2018. I had to squint a couple times to make sure I was reading it right! The year in the URL is easier to read.
- ncmncm 6y agoThe article is an argument against types, in general. The point that characters can be stored in other containers is meaningless: the question is whether, conceptually, a specific sequence of character values distinct from another sequence has compile-time meaning. It does. Therefore, it needs a type. Such a sequence has numerous special characteristics. In particular, element at [i] often has an essential connection to element at [i+1] such that swapping them could turn a valid string to an invalid one. In fact, that an invalid sequence is even possible is another such characteristic.
- arcbyte 6y agoI actually read it as a argument FOR types and against modern languages choice to make the String class a weak proxy for typeless byte arrays. See all the arguments (in this HN comments no less!) for just using utf8 byte arrays as strings. Hes saying semantically there's no difference between arrays and string classes except that with string classes we let you do all kinds of dangerous byte manipulation that we would never dream of with any other type. Moreover, most of the uses for this dangerous access aren't real usages because if you're manipulating strings you're almost certainly actually manipulating code points. So why wouldn't you just use a code point array and give yourself real type safety instead?
- ncmncm 6y agoI did not get that at all. Anyway a code point array would not serve the purpose: most possible sequences of valid code points are not valid strings. A variable-size array of code points is also useful, just as, in C++, a std::vector<char> is useful, but that doesn't make it a string. That C++ std::string<> is wrong for what we now think of as strings is a whole other argument. People once hoped that std::string<wchar_t> or std::string<char32_t> might be the useful string, but they were disappointed. C++ does not have a useful string type at this time, but there is ongoing work on one. It should appear in C++26.
- AnimalMuppet 6y ago> most possible sequences of valid code points are not valid strings. Could you clarify? In what way are they not valid strings?
- 37ef_ced3 6y agoGo's immutable UTF-8 string type is one of the nice things about the language A Go string is almost exactly like this C struct: struct String { uint8_t* addr; ptrdiff_t len; }; The language guarantees you can't modify the bytes in memory range [addr, addr+len) Go's garbage collection makes it simple and natural to have one string alias ("point into", "overlap") part of another string. This works because strings are immutable. Compare this to the nightmare in C++, where substrings require copying or explicit handling The rune (UTF-8) iterator and other facilities make Unicode handling natural in Go In summary, Go's string type is a huge win
- jrimbault 6y agoI'd arguee Go's string type is "somewhat unusable"* since it doesn't enforce the guarantees it says/implies it does. The byte slice it points to is not guaranteed to be valid utf8. * of course to a degree, let's be reasonable, it's usable in a _lot_ of contexts, but I like my types to actually mean something.
- 37ef_ced3 6y agoIn Go, malformed UTF-8 encodings are expected They are handled in a well-defined and graceful manner by all aspects of the language, runtime, and library
- DougBTX 6y agoGo doesn't guarantee any encoding for strings, very deliberately (so that, eg, they can be used to represent file names).
- KMag 6y agoFilesystem paths are not strings. Linux doesn't enforce an encoding. Windows at least didn't used to enforce proper use of conjugate UTF-16 pairs (see WTF-8 encoding). I think OS X does perform UTF-8 normalization, which might include sanity checking and rejecting malformed UTF-8, but I'm not sure. A byte array (or a ref-counted singly-linked list of immutable byte arrays to save space/copying) is a much better representation for a file system path. That doesn't have great interaction with GUIs, but there are other corner cases that are often problematic for GUIs. In high school, one of my friends had a habit of putting games on the school library computers, and renaming them to names with non-printable characters using alt+number pad. (He used 129, IIRC, which isn't assigned a character in CP-1252.) The Windows 95 graphical shell would convert the non-printable characters to spaces for display, but when the librarian tried to delete the games, it would pass the display name to the kernel, which would complain that the presented path didn't exist.
- tyingq 6y agoI can't speak for C++, but for C, the repeated issue is that a null-terminated string has lots of utility routines that are handy for manipulating them. Without 3rd party libraries, plain length-header buffers don't. Hence things like Antirez's sds library, which by nature, is a compromise. I get you can't fundamentally change C now, but a buffer type with a rich manipulation library would have been nice.
- irogers 6y agoString should be an interface/protocol. When I log a message, I want to pass a string. If I have to append large strings for a log message I don't want to run out of memory, I should be able to pass a rope/cord [1]. We've known how to abstract this for forever and should work to optimize our compilers/runtimes accordingly. I'm not aware of a language which has got this right, for example, Java has the ugly CharSequence interface that nobody uses. StringProtocol in Swift (can I implement it?) makes you pay a character tax rather than to just pass a string. Rust/C++ give various non-abstracted types. [1] https://en.wikipedia.org/wiki/Rope_(data_structure) https://en.wikipedia.org/wiki/Rope_(data_structure)
- pdimitar 6y agoErlang/Elixir's iolists which are heavily utilized in Phoenix's templating engine are a rope and are extremely efficient (for a dynamic language). Phoenix's templating is very fast.
- 60secz 6y agoCan't agree more. Java in particular suffers greatly from Object toString with a weak contract and no global String interface. If String were an interface instead of an implementation than any method signature could accept multiple implementations. This allows for really effective type aliases which even support strong typing so if you have a signature with multiple String values you can use the strong types to ensure you don't transpose arguments.
- shadowgovt 6y agoI think the author started from an assertion ("This primary difference between a C++ ‘string’ and ‘vector’ is really just a historical oddity that many programs don’t even need anymore") that highlights an error in the C++ model of strings, not in the way we must think about strings. Contrast NSString in Cocoa (https://developer.apple.com/documentation/foundation/nsstring https://developer.apple.com/documentation/foundation/nsstrin...). The Cocoa string is extremely opaque; it's basically an object. And under the hood, that opacity allows for piles of optimization that are unsafe if the developer is allowed to treat the thing as just a vector of bytes or codepoints. Under the hood, Cocoa does all kinds of fanciness to the memory representation of the string (automatically building and cutting cords, "interning" short strings so that multiple copies of the string are just pointers to the same memory, caching of some transforms under the assumption that if it's needed once, it's often needed again). Taken this way, one can even start to talk about things like "Why does 'indexing' into a string always return a character, instead of, say, a word?" and other questions that are harder to get into if one assumes a string is just 'vector of characters' or 'vector of bytes.'
- BlueTemplar 6y agoToday I learned that Python does interning of shorts strings too : https://news.ycombinator.com/item?id=26097732 https://news.ycombinator.com/item?id=26097732
- deleted 6y ago[deleted]
- hollasch 6y agoCurious. I have to come to exactly the opposite conclusion — that we should drop the idea of a fixed-length character type, and instead _only_ have (Unicode) string types. Actually, I'd prefer something like `std::text` to finally be free of the baggage of "string". Operations on text should work on logical text concepts. For example, something like `someText.firstCharacter()` would have a return type of `text`, with logical length 1. It's _data_ length is variable, since a Unicode character is variable length. So many Unicode-containing string design problems arise because of the stubborn insistence of having an integral character type. I should be able to extract UTF-8, UTF-16 or whatever encoding I want from a `text` value. Something like `c_str()` would be pretty important, but the semantics would be a design problem, not an encoding problem. Any Unicode-encoding string should be able to encode U+0000, so you'd need to figure out how to handle that from `c_str()` (perhaps a substitution ASCII character could be specified to encode embedded nulls). Basically, users should definitely _not_ need to understand the deeper details of Unicode. They shouldn't need to understand and worry about different entities such as code units, code points, graphemes, and the like, though they should be able to extract such encodings on demand.
- shadowgovt 6y agoEssentially, different tools for different applications. "A string is a vector of characters, which happen to each be one byte in length" was more of an artifact of a time where there happened to be representational overlap than some deep truism about proper data structure. Strings intended to be displayed to humans are specialized constructs, much as a "button" or a "file handle" are. A buffer of unstructured bytes is a separate specialized construct, suitable for tasks unrelated to "displaying text to a human."
- lisper 6y agoI fully endorse the general idea here, but this: > `someText.firstCharacter()` would have a return type of `text`, with logical length 1 is a huge mistake. There are operations that make sense on characters that do not make sense on texts whose length happens to be 1. The most obvious of these is inquiring about the numerical value of the unicode code point of a character. Conflating characters and texts-of-length-1 is a mistake of the same order as conflating strings and byte vectors. Python makes this mistake even in version 3. As a result, a function like this: def f(s, n, m): return ord(s[n:m]) will return a value iff m is one more than n. Not good.
- dang 6y agoDiscussed at the time: https://news.ycombinator.com/item?id=6204427 https://news.ycombinator.com/item?id=6204427
- BlueTemplar 6y agoThe author has these followup blogposts : 2013 : https://mortoray.com/2013/11/27/the-string-type-is-broken/ https://mortoray.com/2013/11/27/the-string-type-is-broken/ 2014 : https://mortoray.com/2014/03/17/strings-and-text-are-not-the-same/ https://mortoray.com/2014/03/17/strings-and-text-are-not-the... (See also : https://thehardcorecoder.com/2014/04/15/data-text-and-strings-oh-my/ https://thehardcorecoder.com/2014/04/15/data-text-and-string... ) 2016 : https://mortoray.com/2016/04/28/what-is-the-length-of-a-string-a-tricky-question/ https://mortoray.com/2016/04/28/what-is-the-length-of-a-stri...
- BlueTemplar 6y agoTL;DR : Characters and Strings considered harmful. And he's right, they totally are ! (Also, 'string' can mean an ordered sequence of similar objects of any kind, not just characters.) But (as these discussions also mention) replacing them by much more clearly defined concepts like byte arrays, codepoints, glyphs, grapheme clusters and text fields is only the first step... The big question (these days) is what to do with text, specifically the 'code' kind of text (either programming or markup, and poor separation between 'plain' text and code keeps causing security issues). To start with, even code needs formatting, specifically some way to signal a new line, or it will end up unreadable. Then, code can't be just arbitrary Unicode text, some limits have to apply, because Unicode can get verrrry 'fancy' ! (Arbitrary Unicode is fine in text fields and comments embedded in code.) So, I'm curious, is there any Unicode normalization specifically designed for code ? (If not, why, and which is the closest one ?) I'm thinking of Python (3), which has what seems to be a somewhat arbitrary list of what can and what can't be used as a variable name ? (And the language itself seemingly only uses ASCII, though this shouldn't be a restriction for programming/markup languages !) Also I hear that Julia goes much further than that (with even (La)TeX-like shortcuts for characters that might not be available on some keyboards), what kind of 'normalization' have they adopted ?
- eigenspace 6y agoYes, Julia really lets one get wild with Unicode. There are certain classes of unicode characters that we have marked as invalid for identifiers, some which are used for infix operators, and some which count as modifiers on previously typed characters which is useful for creating new infix operators, e.g. one might define julia> +²(x, y) = x^2 + y^2 +² (generic function with 1 method) such that julia> -2 +² 3 13 If someone doesn't know how to type this, they can just hit the `?` button to open help mode in the repl and then paste it: help?> +² "+²" can be typed by +\^2<tab> search: +² No documentation found. +² is a Function. # 1 method for generic function "+²": [1] +²(x, y) in Main at REPL[65]:1 Note how it says "+²" can be typed by +\^2<tab> Generally speaking we don't have a ton of strict rules on unicode, but it's a community convention that if you have a public facing API that uses unicode, you should provide an alternative unicode-free API. This works pretty well for us, and I think can be quite useful for some mathematical code if you don't overdo it (the above example was not an example of 'responsible' use). I know we have a code formatter, but it doesn't do any unicode normalization. We generally just accept unicode as a first class citizen in code. This tends to cause some programmers to 'clutch their pearls' and act horrified, but in practice it works well. Maybe just because we have a cohesive community though
- BlueTemplar 6y agoAnyone else thinks that we missed an opportunity to make text much simpler to deal with by not increasing the size of a byte from 8 to 32 bits when we moved from 32-bit to 64-bit word length CPUs ? I mean, isn't the 7-bit ASCII text the reason why the byte length was standardized to the next power of two bits ? (With e-mail still supporting non-padded 7-bit ASCII until recently for performance reasons.)
- giardini 6y agoSurprising to a Tcl programmer!8-)) b/c "Everything is a String": https://wiki.tcl-lang.org/page/everything+is+a+string https://wiki.tcl-lang.org/page/everything+is+a+string and "Everything is a Symbol": https://wiki.tcl-lang.org/page/Everything+is+a+Symbol https://wiki.tcl-lang.org/page/Everything+is+a+Symbol
- BlueTemplar 6y agoLooks like that what Tcl means by 'string', the author names 'text' ? What does Tcl mean by 'character' ? See for instance, the author's HTML example : > Combining characters can create an accented version of that symbol, <̧. In text this is clearly a different symbol: it’s a distinct grapheme cluster. The HTML parser doesn’t care about that. It sees code #60 followed by #807 (combining cedilla). It thus sees the opening of an element. However, since it isn’t followed by a valid naming character most parsers just ignore this element (I’m not positive that is correct to do). This is not the case with an accented quote, like "̧. Here the parsers (at least the browsers I tested), let the quote end an attribute and then have a garbage character lying around. https://mortoray.com/2014/03/17/strings-and-text-are-not-the-same/ https://mortoray.com/2014/03/17/strings-and-text-are-not-the... EDIT: Ok, it looks like by 'character', Tcl means what the author (and Unicode ?) calls a 'grapheme cluster' ? https://wiki.tcl-lang.org/page/Characters%2C+glyphs%2C+code%2Dpoints%2C+and+byte%2Dsequences https://wiki.tcl-lang.org/page/Characters%2C+glyphs%2C+code%... https://mortoray.com/2016/04/28/what-is-the-length-of-a-string-a-tricky-question/ https://mortoray.com/2016/04/28/what-is-the-length-of-a-stri...
- quelsolaar 6y agoI think that the problem with text is that the basic operation you want to do is inserts. The way memory works in computer makes that an inherently inefficient operation. I'm a bit fascinated by how bad computer are at text give that that is what we use so much of them for. As a C programmer I think that its not really possible to implement an efficient text processing library, because there is no good universal way to store text. So much depends on the pattern of the processing functions. If you want to avoid allocating new memory and moving a lot of text for each operation, the implementation needs to make speculative choices about how text can best be stored. How you store text depends so much on your access pattern. Do you need to be able to get to a line fast? or know how long the text is? Or insert something? and if so how much? A C style string would for instance be terrible for something like a text editor, because every key press would cause a complete copy of the document to have to be allocated, and then copied over. So maybe a linked list? But you dont want just one character in each link because that trashes the cache right? but then its still slow to just skip forward fast, so maybe an array of pointers to snipets? or maybe a linked list of pointers to snippets? So many possibilities that all impact performance differently depending on what you do with it. When I see higher languages with nice easy to use string functionality, I always consider, the impossible choices that had to be made under the hood.
- xscott 6y agoI think you want a "gap buffer".
- quelsolaar 6y agoA gap buffer is an example of a data structure for text that is optimized for one usage pattern, and performs badly with other patterns. Generalized text structures are hard.