6 ms·
There are way more than 4, but each one makes sense and has a purpose. Trying to conflate them would make things conceptually less clear. Also, somehow people
by hsivonen 9y ago
There are way more than 4, but each one makes sense and has a purpose. Trying to conflate them would make things conceptually less clear.
Also, somehow people seem to object less to Vec<T> vs. &[T] than String vs. &str.
It's great that owned string with contents on the heap is clearly distinguished from a borrowed view into a string. In C, when you see a char*, do you own the string or are you borrowing a string? The type doesn't say. In C++17, however, the distinction is present: std::string vs. std::string_view.
It's also great that Rust, unlike C or C++, distinguishes between strings that are valid UTF-8 and strings that came from a kernel that takes a GIGO position on encodings. Got a std::string in C++? Is it UTF-8 or garbage you got from the kernel? You don't know.
- comex 9y agoThe only thing that bothers me is that Rust’s standard library has no good solution for when you do want to take a GIGO position, such as if you need to interoperate over FFI with C/C++ code that does so. There’s Vec<u8>, but it’s missing a lot of the convenience methods and other functionality available for strings, and its Debug impl (i.e. string representation for debugging) gives you a list of numbers rather than, say, a best-effort UTF-8 interpretation. There’s also OsString on Unix, but that’s OS-specific and also is missing functionality. Unsurprisingly, there’s a third-party crate that has the functionality what I want, and thanks to Cargo that’s almost as good as it being in the standard library… but I still want there to be something in the standard library. :)
- LambdaComplex 9y agoOsString isn't Unix-specific. Or am I misunderstanding what you're saying?
- kibwen 9y agoI think what they're saying is that because Unix doesn't specify any string encoding at all, you can get GIGO by using OsStr on Unix.
- comex 9y agoYes, that's what I meant. Actually, both the Unix variant, "GIGO byte sequence probably UTF-8", and the Windows variant, "GIGO u16 sequence probably UTF-16", have cross-platform use cases. For example, the UTF-16 variant could be used when doing FFI with Java, or when reading Windows filesystems. Therefore, I think ideally Rust would provide both variants as cross-platform types, and OsString/OsStr would just be aliases for one or the other depending on the current platform.
- saghm 9y agoNaive question from someone who's never written FFI code: does std::ffi::CString work here, or is the issue that you need to be able to have null bytes within the string?
- kibwen 9y agoLooking at the docs for CString (https://doc.rust-lang.org/std/ffi/struct.CString.html https://doc.rust-lang.org/std/ffi/struct.CString.html): "An instance of this type is a static guarantee that the underlying bytes contain no interior 0 bytes ("nul characters") and that the final byte is 0 ("nul terminator")." The functions that create CStrings also return a Result where you'll get an Err if your string contains any nul bytes. There's are unsafe variants that don't do these checks, but it's hard to predict what might break if this invariant isn't upheld. There's not much reason to use CString rather than just Vec<u8>.
- kibwen 9y ago> There’s Vec<u8>, but it’s missing a lot of the convenience methods and other functionality available for strings This is where you've lost me, because Vec<u8> is precisely what one should use in Rust for passing around opaque buffers as you describe. I can't even imagine what sort of convenience methods one could implement on strings that deliberately have no specified representation!
- dbaupp 9y agoOne might have strings that have unspecified, but identical, representations, meaning one can do things like 'split' and 'replace' (in general, Vec/&[] doesn't have great tools for operating on subslices, like those two string functions: essentially just 'chunks' and 'windows').
- burntsushi 9y ago> I can't even imagine what sort of convenience methods one could implement on strings that deliberately have no specified representation! A lot. Go's `string` type, for example, is conventionally UTF-8 encoded, but it is not enforced. Consider, for example, how much simpler this code[1] could be if `Vec<u8>` was more convenient to use as a string type. There's a reason why the regex crate provides an API for both `&str` and `&[u8]`, because being able to deal with `&[u8]` as if it were a string is occasionally convenient. Importantly, without this API, ripgrep couldn't feasibly exist! Other examples include file path handling. I need to be able to run globs (via regexes) on them, and the only way I can do that in a way that is zero cost on Unix is to get the bytes from the underlying `&OsStr` (on Windows, I do a lossy decode, which avoids the extra allocation in most places, but still requires the UTF-8 check). In particular, on Unix, this isn't even a matter of performance but rather of correctness, since file paths can contain arbitrary bytes and indeed have no specified representation! (Other than some rules like "no NULs and no /.") It is often very convenient, in practice, to simply assume UTF-8 or at least an ASCII compatible encoding, rather than enforcing it as an invariant. For example, the link above to the ripgrep config parsing assumes the file contains ASCII-compatible text on Unix, and that's it. The file could contain latin-1 or UTF-8, it doesn't matter, and that is required for correctness. (Because file itself could contain file paths which may be arbitrary bytes.) (To be clear, I think the UTF-8 invariant for String/&str was the best choice, certainly. What I'm saying here isn't that one should use Vec<u8> for strings in lieu of String, but rather, that using Vec<u8> as a string can be extremely useful in certain circumstances.) [1] - https://github.com/BurntSushi/ripgrep/blob/7120f3225862f6c718a37a8616debaebd8c3d459/src/config.rs#L77-L103 https://github.com/BurntSushi/ripgrep/blob/7120f3225862f6c71...