4 ms·
> Should byte strings be added to std? > Some folks have expressed a desire for bstr or something like it to be put into the standard library. I’m not sure how
by Fiahil 4y ago
> Should byte strings be added to std?
> Some folks have expressed a desire for bstr or something like it to be put into the standard library. I’m not sure how I feel about wholesale adopting bstr as it is. bstr is somewhat opinionated in that it provides several Unicode operations (like grapheme, word and sentence segmentation), for example, that std has deliberately chosen to leave to the crate ecosystem.
Yes, ok, but could we -at least- have the same Debug impl bstr has ? I'd love to be able to print "human-readable" Vec<u8> :')
- burntsushi 4y agoThat's what the very next paragraph addresses haha. So the Debug impl for Vec<u8> is just the Debug impl for Vec<T>. Doing otherwise means specializing for Vec<u8>, and it's not totally clear to me that it makes sense to do that. Doing it effectively requires assuming that a Vec<u8> everywhere is UTF-8 or close to it. I do mention that we could add a '[u8]::debug_utf8()' method that returns a type with a nice Debug impl for byte strings. Kind of like how we have 'Path::display()', but for the Display impl. But that is kind of annoying in a way that doesn't really apply to Display impls. It's very common to derive(Debug), and if the debug impl is only accessible via a method, then derive(Debug) doesn't work. So then you have to write your Debug impl by hand, which is... annoying. Anyway, point is, it's just not totally straight-forward to bring bstr into std. As I said in the blog, I think the highest value thing that could be brought into std is substring search that works on &[u8].
- dhosek 4y agoI’ve toyed with seeing about adding a feature to finl_unicode to extend or replace the bstr implementations of segmentation, etc. but I don’t need it so I probably won’t. You’re welcome to steal my code though. (And hi from reddit-land!)
- fanf2 4y agoI am curious how bstr relates to OsString. Is the difference that OsString can be WTF16 on Windows?
- burntsushi 4y agoOsStr uses WTF-8 on Windows, and just represents the raw underlying bytes on Unix. Byte strings can be WTF-8. They can be anything. The problem is that there is no real way to (easily) get the underlying WTF-8 bytes of an OsStr on Windows. So there's no free conversion to and from byte strings. I wrote more about this in the bstr docs (and don't miss the link to os_str_bytes): https://docs.rs/bstr/latest/bstr/#file-paths-and-os-strings https://docs.rs/bstr/latest/bstr/#file-paths-and-os-strings I'd be happy to answer more questions if you have them. :-) https://github.com/BurntSushi/bstr/discussions https://github.com/BurntSushi/bstr/discussions
- zozbot234 4y ago> It's very common to derive(Debug), and if the debug impl is only accessible via a method, then derive(Debug) doesn't work. The right way of doing this is to define custom attributes as part of derive(Debug) and derive(Display); the derive mechanism can already do this. There's no need for a wrapper type to be used.
- burntsushi 4y agoYes, it may very well be the case that this is the answer. But that feature is not stable today. It would be great to re-evaluate once that's available.
- IshKebab 4y agoThat shouldn't be the default debug implementation. Plenty of people use `Vec<u8>` to store a list of numbers, not a string.