8 ms·
No thanks. Rust made the right call by distinguishing between owned (String) and borrowed (str) strings. Even C++ is moving in that direction now, by finally
by weirdwitch 9y ago
No thanks. Rust made the right call by distinguishing between owned (String) and borrowed (str) strings. Even C++ is moving in that direction now, by finally adding string_views (https://en.cppreference.com/w/cpp/string/basic_string_view https://en.cppreference.com/w/cpp/string/basic_string_view)
- MBCook 9y agoWhy is one capitalized and one lowercase? I kind of agree with the sibling that the naming is odd. Why not call them String and Slice or String and MutableString? Using different ‘spellings” of the same word seems like it only encourage some confusion.
- steveklabnik 9y agoThe lowercase one is a language primitive, and so is lowercased like all language primitives. The uppercase one is a standard library type, and so is uppercased like all library types.
- juliangoldsmith 9y agostd::String is a struct, and str is a type defined by the language. There's a summary of the differences at [0]. [0]: http://www.ameyalokare.com/rust/2017/10/12/rust-str-vs-String.html http://www.ameyalokare.com/rust/2017/10/12/rust-str-vs-Strin...
- masklinn 9y ago> Why is one capitalized and one lowercase? Because one is a pretty bog-standard struct: https://doc.rust-lang.org/src/alloc/string.rs.html#294-296 https://doc.rust-lang.org/src/alloc/string.rs.html#294-296 while the other is a primitive type: https://doc.rust-lang.org/std/#primitives https://doc.rust-lang.org/std/#primitives Lowercase (i8, u64) or anonymity (`[]`, &, *) means you're dealing with something fundamental to the language. Not just with special compiler/language support (like Result) but way below that, something the language doesn't itself express but has intrinsic knowledge of.
- ComputerGuru 9y agoThat's not the problem. There absolutely should be two different types (and I don't even care that the names are so poor). But half the rust APIs take one type and the other take another (even when the string isn't manipulated or stored in any way, shape, or form). Some interfaces are only implemented for string and others only for &str. Deciding betweene two distinct types &str and &string (not mut &string) for your function's interface is nonsense. It makes no sense to have to _decide_ between which two views of a string that you can read-but-not-manipulate you want to use, and it makes zero sense that they can't unify the types with some simple compiler magic. A constant reference to a string should automatically decompose into a view of that string and that should be that. [edit: as in that view shouldn't be a separate type] Additionally, that dereferencing a string returns a pointer... that makes no sense. That's the kind of nonsense we ran away from in the C++ world. strings are the reason I regret not adopting rust back when as a user of a pre-1.0 language I could have joined in efforts to lobby against this insanity. --- As a sidenote, string_view is so late in coming to the c++ world that it's not even funny. Having a separate std::string with an "implementation-defined" in-memory representation in a world of c strings (char *) is inane beyond belief. (Yes, nulls in strings would still be a problem. But why do your strings have nulls in the first place? That data should probably be a vector of strings or a [vector|array] of uint8_t (even if just typedef'd to unsigned char) and C++ strings should have been mandated utf8, contiguous, and null-terminated. You should be able to compose a zero-copy, read-only, non-owned string from a character array and decompose automatically to it. And don't get me started on the fact that C++ doesn't have sprintf because of the obsession with sticking to the overly verbose and way too complicated streaming operators. Developers end up using c strings with sprintf to format text and then copy it back to a std::string just to work around that stupidity.
- weirdwitch 9y agoAnything implemented for &str is automatically implemented for String, because String implements Deref<Target=str>. Most useful "String" methods are actually &str methods that you get access to through that deref trait. Dereferencing a String doesn't return a raw pointer, I'm not sure where you got that idea.
- 9y ago
- simfoo 9y agostr and String? That naming is horrible. string_view is a much better choice
- tomjakubowski 9y agoIt's certainly confusing at first, but the terse naming of "str" is appreciated once you've internalized the difference. "str" is, in most projects I've seen, appears more often in source than "String".
- masklinn 9y ago1. string_view makes for much more verbose code when it's the more common version. 2. string_view makes no sense when many str are not actually views into strings but either static data or "cast" bytes. 3. string_view makes no sense when str is the basic builtin type. Renaming String to something like StrBuf might have made sense, but it's not like it would have been any clearer, people get confused by Path/PathBuf all the same if not more so.
- weirdwitch 9y agoI completely agree. Path/PathBuf have better names, but the distinction still isn't immediately obvious. String/str is the cause of so much confusion.
- steveklabnik 9y agoI often joke that str/StrBuf is my one wish for a Rust 2.0, but 2.0 will never happen. This confusion is why we reorganized the book to talk about String and &str very early on, and use them to teach ownership and borrowing.
- Retra 9y agoIt took me like 2 seconds to learn it and haven't been confused since. So... maybe I'm just a grade A genius.
- andrewmcwatters 9y agoCould you explain to a non-Rust user why this is a good thing? I'm not familiar with the terms "owned" and "borrowed" in terms of strings. I take it this refers to strings instantiated by a piece of code and managed by a user or vendor code vs strings passed around to be read (pass by copy)?
- masklinn 9y ago> I'm not familiar with the terms "owned" and "borrowed" in terms of strings. A string (general) is fundamentally a bunch of bytes in memory. In Rust, that's implemented as a contiguous buffer of UTF8 code units composing valid unicode data. Now because Rust aims to be a systems-oriented language, these bytes must have one[0] thing which is fundamentally responsible for them, that's the owner. If the owner goes away, so do the bytes. "Owned" qualifies the owner of those bytes. In Rust, it's String: String maintains a "strong" reference to a bunch of bytes in memory which form a valid string, and if the String disappear so does the data associated with it (it's deallocated). Borrowed by comparison is something which holds a "weak" reference to the same buffer, it knows they're there, but when the "borrowing" structure is destroyed nothing happens to the data it refers to, because it was just borrowing it. That's what `&str` is. Note that `&str` can borrow from a String (that's a common case), but it can also borrow from static data in the binary (all string literals are in that case) or from just a random bytes buffer (using str::from_utf8[1]). > Could you explain to a non-Rust user why this is a good thing? It doesn't matter for high-level languages like Java or Python[2] but it matters a lot for lower-level languages like C, C++ or Rust, because the owner of a piece of data is whoever's supposed to deallocate it (and whoever's allowed to reallocate it to expand it). When there's no difference between owner and borrower (e.g. C's char * ) the complete onus of tracking who's responsible for what is on the developer and failure brings for them to do so generates memory unsafety (dangling pointers, double-free, use-after-free, …). And developers being humans, they mess up regularly. Making a very clear distinction between owned and borrowed types firstly helps the developer: they know they must free through the owner but mustn't — and normally can't — free through the borrower, that's what C++ is adding; and secondly can — with some additional constraints on the developer — have the language manage all that on its own, the latter being what Rust does (C++ will do the freeing part but you can still have extant borrows and so it's not memory-safe, just significantly more helpful than C). [0] unless you're in a case where it's unclear who should be responsible and you just go "everyone!" and use reference-counting to just punt [1] https://doc.rust-lang.org/std/str/fn.from_utf8.html https://doc.rust-lang.org/std/str/fn.from_utf8.html [2] generally, it does matter in the problem of substrings and whether substrings copy or point to the base data, the former is more expensive (you allocate on each substring operation) but the latter can maintain gargantuan amount of data "live" and prevent their collection