4 ms·
I've been using this library in production to handle doing operations on domain names and it's been incredible. It's one of those things that's so easy to use i
by vlmutolo 4y ago
I've been using this library in production to handle doing operations on domain names and it's been incredible. It's one of those things that's so easy to use it almost starts to seem simple. Like of course we need a library that looks just like this. It's obvious in hindsight, which speaks to great design.
It's especially helpful that the library doesn't require you to opt into its own dedicated types, and instead defines extension methods on existing types.
Thanks, Andrew!
- burntsushi 4y ago> It's especially helpful that the library doesn't require you to opt into its own dedicated types, and instead defines extension methods on existing types. Fun fact: bstr 0.1 went the route of defining its own dedicated types! See: https://docs.rs/bstr/0.1.4/bstr/ https://docs.rs/bstr/0.1.4/bstr/ But it did indeed quickly prove to be pretty annoying. Because you still really want to use &[u8] in places because it's so ubiquitous. But to get access to the byte string methods, you had to explicitly convert it to another type. The reason why I went that route initially was so you'd always get the good Debug impl. But it ended up not being worth it sadly. This issue discusses it a bit more: https://github.com/BurntSushi/bstr/issues/5 https://github.com/BurntSushi/bstr/issues/5
- epilys 4y agoThis should be in the standard library, but maybe I'm biased. In implementing ascii/utf8 plain text protocols like IMAP/SMTP just like the grep/ripgrep example in the article, I've had to limit myself to u8 slices just because the occasional byte might make a grapheme invalid.
- burntsushi 4y agoYeah people say this enough that I added a section about it to the blog: https://blog.burntsushi.net/bstr/#should-byte-strings-be-added-to-std https://blog.burntsushi.net/bstr/#should-byte-strings-be-add...
- epilys 4y agoThe needle/haystack thing is the crucial part as you say in the blog. Since grapheme, word separation and sentence segmentation rely on rules that are updated on each unicode version it indeed would be out of place in the standard library. Ideally I'd want them (libicu) on the system level, e.g. standardised in POSIX in libc so that language developers shouldn't need to worry about them.