7 ms·
This feature of strings is cool! It sounds like the end of the Unicode encoding mess that most languages drag the programmer into: "One of the truly clever des
by Normati 12y ago
This feature of strings is cool! It sounds like the end of the Unicode encoding mess that most languages drag the programmer into:
"One of the truly clever design choices for Swift's String is the internal use of encoding-independent Unicode characters, with exposed "views" to specific encodings:
A collection of UTF-8 code units (accessed with the string’s utf8 property)
A collection of UTF-16 code units (accessed with the string’s utf16 property)
A collection of 21-bit Unicode scalar values, equivalent to the string’s UTF-32 encoding form (accessed with the string's unicodeScalars property)"
- richardwhiuk 12y agoIt sounds like a fairly terrible tradeoff actually. Presumably one of those encodings is the underlying representation, and accessing it in any other representation is going to cause a horrible hit to performance as it thunks between the two. In practice it feels much better to define a encoding for the platform (say UTF 16 for Java / Windows - though personally I feel UTF-8 is a better choice) and then have a [string encodeIn:UTF16] method if you want something different. Ideally, the encoding would be specified in the type - e.g. String<UTF16> rather than arbitrary unspecified byte arrays.
- masklinn 12y ago> Presumably one of those encodings is the underlying representation, and accessing it in any other representation is going to cause a horrible hit to performance as it thunks between the two. Which does not usually matter, you're accessing it in a specific representation because you need it in that representation, usually for IO. The "horrible hit" is one you'll have to eat either way. And if you're baking the implementation details of your internal strings into your IO… god help your soul. That aside, there's not much of a horrible hit unless you're preallocating the whole output string every time. Swift has iterators/iterables built in and I may be mistaken but I believe Swift does the sane thing and exposes noalloc iterable views, you're paying for some bit-twiddling (for the transcoding itself) and stack-allocated int8/int16/int32. Not sure how good Swift's compiler is, but I know Rust's can turn such iterations into the equivalent of the corresponding C loop, there's little to no overhead. > In practice it feels much better to define a encoding for the platform Why? What does that give you, aside from exposing broken implementation details as the type's public interface and having unfixable string for a decade (see: Java and everything Microsoft, because they exposed strings as being O(1)-indexed UCS2 code units early on). > Ideally, the encoding would be specified in the type - e.g. String<UTF16> That's crazy talk, why would you encode the implementation detail of the string's internal encoding in the type interface? I can think of a hundred things I'd put there, but the internal encoding?
- dilap 12y agoHaven't used Swift of its strings yet, but I have used string/unicode in python, NSString in Cococa, and string in Go. The only one that hasn't bit me in the ass is string in Go; the approach they take is a sequence of bytes, utf8 encoded by convention (which is easy to follow); there are methods to work with unicode code points when you need to, otherwise it's just bytes. It's simple, well-defined, non-magic, and doesn't screw you over.
- bpicolo 12y agoPython 3 uses utf8-as-a-default as well. Not to mention you can easily specify the encoding as well in python 2. Never bites me.
- dilap 12y agoWish I had documented my exact frustrations, but it wasn't not being able to choose the encoding...something more of the flavor of: Most of the time I just trying to shuffle bytes from one place to another, and didn't really care about the contents of the bytes, but I hit numerous bug because of exceptions due to things like "external data wasn't valid utf8" (NOT helpful to throw an exception here), and bugs caused by 'str' vs 'unicode' confusion (made more tricky by lack of static type system). The one time I did actually need to do some calculation involving unicode I got burned by behavior that didn't match the documentation, because my bullshit package-manager-provided variant of python had some insane compile time option that made it pretend utf16 code units were the same thing as actual utf code points, to which I can only say "ha ha ha...fuck those guys". Can't speak to python 3, haven't used it, probably never will :)