11 ms·
string is just an immutable []byte. It's actually one of my favorite things about Go that strings can contain invalid utf-8, so you don't end up with the Rust m
by assbuttbuttass 1y ago
string is just an immutable []byte. It's actually one of my favorite things about Go that strings can contain invalid utf-8, so you don't end up with the Rust mess of String vs OSString vs PathBuf vs Vec<u8>. It's all just string
- zozbot234 1y agoRust &str and String are specifically intended for UTF-8 valid text. If you're working with arbitrary byte sequences, that's what &[u8] and Vec<u8> are for in Rust. It's not a "mess", it's just different from what Golang does.
- gf000 1y agoIf anything that will make Rust programs likely to be correct under any strange text input, while Go might just handle the happy path of ASCII inputs. Stuff like this matters a great deal on the standard library level.
- deleted 1y ago[deleted]
- maxdamantus 1y agoIt's never been clear to me where such a type is actually useful. In what cases do you really need to restrict it to valid UTF-8? You should always be able to iterate the code points of a string, whether or not it's valid Unicode. The iterator can either silently replace any errors with replacement characters, or denote the errors by returning eg, `Result<char, Utf8Error>`, depending on the use case. All languages that have tried restricting Unicode afaik have ended up adding workarounds for the fact that real world "text" sometimes has encoding errors and it's often better to just preserve the errors instead of corrupting the data through replacement characters, or just refusing to accept some inputs and crashing the program. In Rust there's bstr/ByteStr (currently being added to std), awkward having to decide which string type to use. In Python there's PEP-383/"surrogateescape", which works because Python strings are not guaranteed valid (they're potentially ill-formed UTF-32 sequences, with a range restriction). Awkward figuring out when to actually use it. In Raku there's UTF8-C8, which is probably the weirdest workaround of all (left as an exercise for the reader to try to understand .. oh, and it also interferes with valid Unicode that's not normalized, because that's another stupid restriction). Meanwhile the Unicode standard itself specifies Unicode strings as being sequences of code units [0][1], so Go is one of the few modern languages that actually implements Unicode (8-bit) strings. Note that at least two out of the three inventors of Go also basically invented UTF-8. [0] https://www.unicode.org/versions/Unicode16.0.0/core-spec/chapter-3/#G32765 https://www.unicode.org/versions/Unicode16.0.0/core-spec/cha... > Unicode string: A code unit sequence containing code units of a particular Unicode encoding form. [1] https://www.unicode.org/versions/Unicode16.0.0/core-spec/chapter-3/#G32860 https://www.unicode.org/versions/Unicode16.0.0/core-spec/cha... > Unicode strings need not contain well-formed code unit sequences under all conditions. This is equivalent to saying that a particular Unicode string need not be in a Unicode encoding form.
- xyzzyz 1y agoThe way Rust handles this is perfectly fine. String type promises its contents are valid UTF-8. When you create it from array of bytes, you have three options: 1) ::from_utf8, which will force you to handle invalid UTF-8 error, 2) ::from_utf8_lossy, which will replace invalid code points with replacement character code point, and 3) from_utf8_unchecked, which will not do the validity check and is explicitly marked as unsafe.
- maxdamantus 1y agoBut there's no option to just construct the string with the invalid bytes. 3) is not for this purpose; it is for when you already know that it is valid. If you use 3) to create a &str/String from invalid bytes, you can't safely use that string as the standard library is unfortunately designed around the assumption that only valid UTF-8 is stored. https://doc.rust-lang.org/std/primitive.str.html#invariant https://doc.rust-lang.org/std/primitive.str.html#invariant > Constructing a non-UTF-8 string slice is not immediate undefined behavior, but any function called on a string slice may assume that it is valid UTF-8, which means that a non-UTF-8 string slice can lead to undefined behavior down the road.
- adastra22 1y agoI don’t understand this complaint. (3) sounds like exactly what you are asking for. And yes, doing unsafe thing is unsafe.
- maxdamantus 1y ago> I don’t understand this complaint. (3) sounds like exactly what you are asking for. And yes, doing unsafe thing is unsafe You're meant to use `unsafe` as a way of limiting the scope of reasoning about safety. Once you construct a `&str` using `from_utf8_unchecked`, you can't safely pass it to any other function without looking at its code and reasoning about whether it's still safe. Also see the actual documentation: https://doc.rust-lang.org/std/primitive.str.html#method.from_utf8_unchecked https://doc.rust-lang.org/std/primitive.str.html#method.from... > Safety: The bytes passed in must be valid UTF-8.