6 ms·
When I was working on one of our SDKs (pre Swift 2.0), I found it rather maddening that the Swift string class had native support for emoji, but developers were
by tigeba 11y ago
When I was working on one of our SDKs (pre Swift 2.0), I found it rather maddening that the Swift string class had native support for emoji, but developers were left to create their own implementation of very basic features like indexOf, length, subString, etc.
- pilif 11y ago... until you realise these "very basic" features are not very basic at all and pretty much depend on your current use-case. What length do you mean? Byte length? Character length? Code Point length? In case of indexOf, how do you handle surrogate pairs? does indexOf('ä') only find 'ä' (LATIN SMALL LETTER A WITH DIAERESIS) or also 'ä' (LATIN SMALL LETTER A followed by COMBINING DIAERESIS)? The good thing about the Swift API is that it gives you all the building blocks needed to actually having a chance at getting this right. Many other languages sweep those things under a rug and you're screwed or you'll have a much, much harder job to get it right if you need to.
- cookiecaper 11y agoI think there's a pretty strong generalized use case for things like length (count the number of characters in a string) and indexOf (return the position of a character in a string). No reason to punish the 99% and force them to write their own implementations of these very conventional string operations to accommodate the 1% that means something not normally meant.
- mikeash 11y agoOnce you decide what a "character" means, these operations are easy. You want the count of grapheme clusters? string.characters.count. You want the index of a particular UTF-16 code unit? string.utf16.indexOf(codeUnit). If you need that index in the form of the number of code units from the start of the string, string.utf16.startIndex.distanceTo(index).
- anon1385 11y agoBased on Github searches I've done in the past a lot of code calling length on NSString was broken. Those people probably thought they were in the 99% who just needed the convenient 'normal' length. People not realising that length isn't giving them what they want is one of the things the Swift API aims to solve. Giving the wrong choice for most situations (and number of UTF-16 code units isn't what people want most of the time) the easy convenient name is just a recipe for broken code. It's bad API design, or at least unfortunate historical accident, and it's good to improve things when there is an opportunity to do so.
- cookiecaper 11y agoIt's not an improvement to exclude functions that you know everyone wants from the API. It's just going to result in more mess, because we all know most people are going to search for "swift string length function" and copy and paste direct from Stack Overflow without more than skimming it. In practice, it just makes the behavior messy and non-standard, ultimately meaning Swift applications are more difficult and/or annoying to debug and maintain. The correct way to solve this would be to design the API so that the distinction between what you're getting and what you want is clear, and the place to go to get what you actually want is also clear. Including nothing is an admission that they couldn't do this.
- anon1385 11y ago>It's not an improvement to exclude functions that you know everyone wants from the API. Which functions have they excluded? You mentioned counting characters, but you haven't actually specified which definition of characters you want it to use. >The correct way to solve this would be to design the API so that the distinction between what you're getting and what you want is clear, and the place to go to get what you actually want is also clear This is what they have done.
- cookiecaper 11y ago>Which functions have they excluded? You mentioned counting characters, but you haven't actually specified which definition of characters you want it to use. The typical human definition of characters. Humans usually don't think about bits or bytes and they shouldn't have to. A character is a single independent glyph, a separate unit that would be taught to a human whilst learning to write, regardless of the internal representation in the computer. If I need something else, I should ask for something else. Hypothetical API calls that may address this: "String".length() "String".lengthInBytes() "String".lengthInCodePoints() This would provide standard implementations for these common functions (circumventing the issue of copying a random chunk of code from SO and all its attendant problems), make it obvious that there is a difference to anyone browsing the docs and/or using autocomplete, and make it easy to select and use the one you actually want in a particular situation. This type of discoverability is an important component in a usable API.
- zw 11y agoBut the definition of "character" is not generalized, therefore the use case cannot be generalized.
- coldtea 11y ago>... until you realise these "very basic" features are not very basic at all and pretty much depend on your current use-case. What length do you mean? Byte length? Character length? Code Point length? Still VERY basic, and all should be provided by the standard library. ("basic" as: very basic and frequent needs. They are of course quite complicated to write. Which is even more of a reason to have them written for developers in the standard library).
- Lx1oG-AWb6h_ZG0 11y ago... This entire article is about how those operations are basic at all, and will cause bugs if you're not careful.