20 ms·
They can put it on the standard library or the core language or wherever they want, but they absolutely need to provide good string handling. And regardless of
by diegocg 5y ago
They can put it on the standard library or the core language or wherever they want, but they absolutely need to provide good string handling.
And regardless of where they put the code, this is something that needs to be done by the core team. Otherwise you will end with too many string libraries, all of them trying to solve a particular problem and doing bad at everything else, with bad documentation and different API styles and making difficult to mix code that uses two different libraries.
- AndyKelley 5y agoA lot of people reach for string handling when the actual correct thing to do is intentionally avoid string handling, and only handle strings as opaque encoded UTF-8 bytes, that cannot be reasoned about in terms of human language. I would even argue that having string handling in a standard library (or language) has the potential to cause a net increase in bugs, because of people thinking they are handling strings when actually they are just screwing around with codepoints. Go's string handling is completely broken, for example. As a result of strings in the language, Go programs tend to be more broken than C programs in terms of string handling.
- gautamcgoel 5y agoCan you elaborate on why you think Go's string handling is broken?
- tdeck 5y ago> the actual correct thing to do is intentionally avoid string handling, That sounds nice and all, but wr have 50+ years of protocols and formats and APIs built up around strings. Unless you're just writing code to run on a small microcontroller, you need to be able to parse and generate strings. So its going to be pretty frustrating not to have good support for them, or to have every codebase use its own libraries and idioms for working with strings.
- AndyKelley 5y agoProtocols and formats absolutely should not require decoding strings. I think you are mistaken. Can you name any well-established protocol or format that does not treat strings as opaque encoded bytes? Edit: so far these examples have been given: * HTTP: wrong. the spec does not tell you to decode any strings * CSV: wrong. the spec does not tell you to decode any strings, nor is it necessary to have any unicode awareness in order to properly read and parse the data or deal with the delimiters. Additional request: if you attempt to provide a counter-example, please also point to the place in the spec where it tells you to decode a string.
- adgjlsfhk1 5y agoCSV?
- ironmagma 5y agoHTTP for one. It does something like assume ASCII-like until it encounters odd looking bytes and/or a meta encoding tag.
- wyldfire 5y agoI think HTML tags like 'meta' are merely payload to HTTP, right? Presumably a markup language doesn't fit Andrew's criteria for "protocol"?
- 10000truths 5y agoProtobuf string fields are expected to be in UTF-8. Although I'd expect a sane implementation of a protobuf decoder to throw an exception or otherwise indicate an error if it receives a malformed protobuf encoding.
- haberman 5y agoProtobuf decoders are expected to validate UTF-8 strings for syntax="proto3" files, but not syntax="proto2". The behavior diverges mostly for historical reasons. This is a validation pass only and it doesn't make any meaning of the code points, except to validate that none of them are surrogate code points (disallowed in UTF-8).
- _rend 5y agoIMO, Swift is a language that gets it right, by exposing an interface which feels to very closely match how "humans" understand text — the default String interface is a collection of grapheme clusters, while `.utf8View`, `.utf16View`, and `.unicodeScalarView` are optional views which expose data explicitly encoded as needed. Swift is pretty pedantically strict about Unicode correctness, and avoids some pitfalls which lead to incorrect handling (e.g., integer-based indexing). This means that the traditional way many developers might be used to interacting with strings is cumbersome and annoying — but once you get past the initial hurdle*, most operations are actually (1) really easy, especially when expressed in terms of generic Sequence/Collection operations, and (2) much more difficult to get wrong. *I think the largest part of that hurdle is overcoming what you may have gotten used to from other languages, i.e., treating strings as an array of "characters", for some language/library definition of "character" (whether bytes, code points, etc.). It's relatively rare that you actually care about indexing into an arbitrary spot in a string: instead, combinations of slicing operations (including `prefix(_:)`, `dropFirst(_:)`, `take(while:)`, etc.) and generic Collection operations will get you what you want. Things like `.reversed()`, `.sorted(), and `.shuffled()` all work trivially correctly too (since you're not operating on a bag of bytes), and it's exceedingly rare that user input will confound the operations you might need to perform. (Exception: operations like case folding and collection, which are locale-specific, need special handling through a framework like Foundation.) To be clear: not everything is sunshine and roses, but an amazing amount of functionality "falls out" of basic protocol conformances on String, and its exposure as a Collection of grapheme clusters. Given a specific string manipulation task, I'd be happy to provide an example of what it might look like in Swift!
- rudedogg 5y ago> Given a specific string manipulation task, I'd be happy to provide an example of what it might look like in Swift! How would you safely get the nth index of a string, clamped to the valid indexes? So if the nth index is out-of-bounds you get the first/last index instead?
- coldtea 5y agoOr you get an Optional and you have to check if there's an error or you go the character within the range? Sounds perfect safe to me...
- pphysch 5y agoYeah strings are ugly and programmatically impure -- they are an extension of human language after all -- but virtually all human facing apps use tons of them for obvious reasons. Same goes for regex. IMO they are a great example of how Golang is a pragmatic rather than "clever" or "pure" programming language.
- skybrian 5y agoI’m not aware of any significant breakage due to Go’s handling of strings. Where would I read more about this?
- noisem4ker 5y agoGo presents strings as slices of bytes (chars). As long as it's ASCII, it's fine. However, when dealing with multi-byte UTF-8 characters (codepoints), the proper unit is the rune. So, before attempting to measure a string's length, or read its n-th character, one must remember and access the string as a slice of runes. s := "naïve" // bad fmt.Println(len(s)) // 6 fmt.Println(string(s[2])) // Ã // good r := []rune(s) fmt.Println(len(r)) // 5 fmt.Println(string(r[2])) // ï https://go.dev/play/p/YbMo49wU7vu https://go.dev/play/p/YbMo49wU7vu
- skybrian 5y agoYes, but I think most Go developers are aware of Go's quirks, so I'm wondering what bugs happen despite that awareness.
- turminal 5y agoBy not treating strings specially, you get people screwing around with individual bytes inside codepoints which leads to at least as many bugs.
- Spex_guy 5y agoThe UTF-8 encoding is designed so that this is usually not a problem. If you do a search in a utf-8 encoded byte array for an ascii character, for example, you can never get a false positive. Compound UTF-8 characters always have the most significant bit set of each component byte, and ascii characters always have it unset. Additionally, treating the string as an array of unicode codepoints doesn't solve the problem -- now you have people screwing around with individual codepoints inside grapheme clusters :P
- turminal 5y ago> Additionally, treating the string as an array of unicode codepoints I suggested no such thing. > individual codepoints inside grapheme clusters That's less severe than invalid codepoints. Perhaps the whole thing whichever way it is represented should not be mutable given that there's no way to make it mutable in a sensible way?
- wyldfire 5y ago> A lot of people reach for string handling when the actual correct thing to do is intentionally avoid string handling, and only handle strings as opaque encoded UTF-8 bytes, that cannot be reasoned about in terms of human language. I have lost count of how many times I have wanted to find substrings, transform cases, catenate strings, find patterns, substitute patterns. I'd be happy to do that in a language that didn't permit me to index the underlyinc characters or the bytes. Keep 'em opaque, sure. But I think it would be a mistake for a language not to have an idiom with a favorite library to perform these operations on encoded text. If resolving library dependencies is easy enough, then it doesn't need to be "standard" but it should be "the defacto standard." And if it turns out the defacto standard stagnates and doesn't keep up with the needs of developers, a new one can come take its place.
- jstimpfle 5y agoMaybe Java or Python are better choices for your problems. The point of a systems programming language is to allow implementation of specific solutions to specific problems. There are many, many ways to do the things you mention here (and that starts with the strings' storage format and allocation strategy) , so it's the right choice to not include _anything_ like that in the core language.
- fastball 5y agoI don't think I've ever had a problem with Python3's string handling, which is very robust.
- bvrmn 5y ago>>> "ñ"[0] 'n' >>> "ñ"[1] '̃'
- fastball 5y agoHuh, which version is that? Python 3.9 on my system: >>> "ñ"[0] 'ñ' >>> "ñ"[1] IndexError: string index out of range Which is what I would expect.
- bvrmn 5y agoPython 3.9.9 (main, Nov 20 2021, 21:30:06) [GCC 11.1.0] on linux
- fastball 5y agoEDIT: I can't seem to replicate this on my Ubuntu system. Strange. Ah, interesting, I suppose then there is a difference in the way Linux handles strings? Didn't realize that, very unfortunate if true. I am running MacOS.
- bvrmn 5y agoActually browser normalized string and it works for me also. Here is original byte sequence: >>> b'n\xcc\x83'.decode() 'ñ' >>> b'n\xcc\x83'.decode()[0] 'n' >>> b'n\xcc\x83'.decode()[1] '̃' But I agree, it's rare case when you need to deal with non-normalized data.
- fiedzia 5y agothis is normalization issue, not version issue import unicodedata list(unicodedata.normalize('NFD', "ñ")) >> ['n', '̃'] list(unicodedata.normalize('NFC', "ñ")) >> ['ñ'] both are correct, the issue is that unicode allows accented letters to be written as _accented_letter_ or _letter_, _accent_. The idea of "character" in uncicode is not very useful, most of the time you will want graphemes, not codepoints. User-friendliness wise, this is what Python should use (another rant - strings should not have length method, they should have byte_length, codepoint_length and grapheme_length).
- svnpenn 5y ago> Go's string handling is completely broken, for example. Classic Andy, shitting on other language with no references or examples. Go has some of the best string handling I've used. Seamless byte, rune, string conversion. Simple iterating and slicing. Plus helpful tools like strings.Builder and strconv.AppendInt. while Zig has nothing.
- deleted 5y ago[deleted]
- coldtea 5y agoIt seems like every language must have some dumb decisions made because the authors don't care/think they know better/are stubborn/etc or because it painted itself into a hole from the start, without which it would be close to perfect for its use cases. Python has the GIL. Go itself has a few such cases (usually revolving around "NIH" and misguided simplicity). Zig has the prejudice about proper string handling.
- bvrmn 5y ago> Python has the GIL. You are welcome to show fast single-core python interpreter without GIL. Coreteam will gladly accept your patches.
- kzemek 5y agoThat's how I understand "painted itself into a hole from the start" from the GP - removing GIL at this point is very hard because virtually all of Python language and code has been created in a world with the GIL.
- bvrmn 5y agoIt's nice to have atomic guarantees of GILed implementation. For me current CPython semantic is a golden middle: you have pretty fast interpreter without race conditions (for huge part of user code). Due to GIL. I always like to ask for what kind of task one needs GILless python? CPU bound? If you are using native python code for CPU bound tasks you are already in a bad place. C extensions can release GIL. For example numpy. What else?
- eyelidlessness 5y agoAs a counterpoint from primarily experience with higher level languages, not being able to distinguish strings from just any blob of bytes is something that’s consistently put me off lower level languages. Sure you can just treat it as a blob of bytes but… you can’t do anything with it with certainty. That seems to be the class of bugs you’re describing? But this isn’t an issue in languages where a string is a distinct type (even if dynamic).