4 ms·
I've never personally understood the exaggeration people use when talking about string types in Haskell. You say "zillions", but there's only really five. Five
by tene 10y ago
I've never personally understood the exaggeration people use when talking about string types in Haskell. You say "zillions", but there's only really five. Five is the most disappointing "zillions" I've ever encountered.
There's strict and lazy ByteString, which you should use whenever you're working with buffers of bytes that don't have any associated encoding and aren't text.
There's strict and lazy Text, which you use for human-readable text that's decoded from some specific encoding, like ASCII or Unicode.
There's String, which you use when interacting with an API that requires use of String, like anything in the language standard.
It's certainly some amount of complexity, but it's not anywhere near as complex as I keep seeing claimed, unless I'm missing something. It's certainly ugly that we still have to deal with String, and that's a notable wart on the language, but I rarely even think about it.
One guess I've had about part of the confusion is that ByteString has the word "String" in it, which might make people assume that it's for text. A better name would be "Buffer" or "Bytes". I've also speculated that maybe part of the complexity is having to explicitly encode and decode to get Text, instead of the implicit coercion you get in other languages that freely mix bytes and text.
If you want a greatly simplified way to unify string conversion, you can use string-conv. https://hackage.haskell.org/package/string-conv-0.1/docs/Data-String-Conv.html https://hackage.haskell.org/package/string-conv-0.1/docs/Dat...
Edited to add: it's been pointed out to me that "zillions" might not have actually been a claim that there's enough string types actually keep track of them, but is probably just hyperbole to express some frustration. If that's the case, my apologies for my failure at reading comprehension.
- bunderbunder 10y agoThe optimal number of core string types for a language to have is one. Python got to two, and the community decided that was enough of a PITA to merit a breaking change in the language in order to get back down to one. C++ also has two, but we put up with it because it's C++ so really we're just thankful it's not three or four. Five is so far beyond the pale that it is absolutely reasonable to hyperbolize it as "zillions".
- tene 10y agoByteString isn't really a "string" type; it's a byte buffer. It's not about text. I think that it's entirely reasonable to have a dedicated "Text" type that works with decoded human text, separate from buffers of bytes. I also think that it's entirely reasonable to have lazy and strict variants of core data types; they have very different time/space tradeoffs. It's definitely frustrating that the core language definition has a data type "String" that's a pretty bad choice for almost any application (it's a linked list of characters), and that's a legitimate problem, but I think an accurate characterization of the real problem with string data types in Haskell is much more like "You should be able to use Text everywhere to deal with decoded unicode text, but for unfortunate legacy reasons you have to use String to interact with large parts of the standard library". Treating all bytes as if they happen to accidentally represent utf-8 encoded unicode text would be a big mistake; there's significant advantages to representing text and byte buffers separately. They are very different things. Choosing to support only strict or only lazy handling of byte buffers or text would be quite unfortunate; there are significant advantages to both for different algorithms and use cases, and encoding the difference in the type system seems entirely reasonable to me. I'm curious which of these you disagree with. Would you prefer that Text and ByteString be merged into one data type, so that the compiler doesn't consider it a mistake to treat arbitrary bytes as if it were text without specifying any encoding? Would you prefer that Text and/or ByteString discard support for lazy representation, or for strict representation?
- bennofs 10y agoDoesn't python3 have both bytes and str, the equavilent of ByteString and Text?
- twic 10y agoRust has six, plus the Path types: http://www.suspectsemantics.com/blog/2016/03/27/string-types-in-rust/ http://www.suspectsemantics.com/blog/2016/03/27/string-types... I would say it's actually justified in that.
- kibwen 10y agoThose aren't string types, they're buffers for platform-dependent system interoperation. When used, the goal is to convert them to String as soon as possible. If one considers those to be string types, then Vec<u8>, Vec<16>, and Vec<32> would be considered string types as well (in addition to both fixed-size arrays and unsafe pointers to the same).