4 ms·
What surprises me most is his insistence on a non-incremental update. To this day, I have no idea what the logic behind "no intermediate steps" is. Codebases ar
by generationP 7y ago
What surprises me most is his insistence on a non-incremental update. To this day, I have no idea what the logic behind "no intermediate steps" is. Codebases are written in Python that would dwarf Margaret Hamilton's famous stack. Why should everyone have to adapt the string functions to unicode and get rid of different-type comparisons and change syntax in one go, without a good way to sound out bugs inbetween?
- riazrizvi 7y agoRegarding string, the breaking change was necessary because it was a double-duty type, sometimes acting as a byte-array and other times acting as a string. Meaning some of its functions, like string.length(), gave a value that only made sense for string-as-byte-array, but not as string-as-simple-list-of-characters. More detail on Stackoverflow if you want (https://stackoverflow.com/questions/5471158/typeerror-str-does-not-support-the-buffer-interface/34688434#34688434 https://stackoverflow.com/questions/5471158/typeerror-str-do...). Anyhow the double-duty type needed to be disambiguated into two separate types, a breaking change.
- papln 7y agostring.length() still doesn't make sense in Python 3 There's nothing "simple" about "list of characters" in Unicode. Discussed at length at: https://news.ycombinator.com/item?id=18154667 https://news.ycombinator.com/item?id=18154667 Here's how a Unicode export with 20 years of experience in Unicode explained it: https://blog.golang.org/strings https://blog.golang.org/strings > Some people think Go strings are always UTF-8, but they are not: only string literals are UTF-8. As we showed in the previous section, string values can contain arbitrary bytes; as we showed in this one, string literals always contain UTF-8 text as long as they have no byte-level escapes. > In fact, the definition of "character" is ambiguous and it would be a mistake to try to resolve the ambiguity by defining that strings are made of characters. > [Exercise: Put an invalid UTF-8 byte sequence into the string. (How?) What happens to the iterations of the loop?] If you aren't in Unicode, then "character" is the same as "byte" so splitting the types gets you nothing of vlaue. Assuming you are not implementing a low-level Unicode library, what's your use case for "list of Unicode characters" ?
- microcolonel 7y agoEven in Unicode-aware applications, it is generally not that useful to segment just by codepoint, you more often really want to segment by extended grapheme cluster.
- takeda 7y agoIn python 2.7 they also introduced "bytes" although it was just another alias to "str". I wonder if they would make python 2.7 distinguish between them and add an option (like another import from __future__) to throw errors when these types were mixed. This would help a lot in preparing code incrementally to work on python3. In the mean time, mypy in python 2 mode doesn't warn about str <-> bytes either :/
- jhall1468 7y agoI mean, it seems like you are suggesting 3 incremental non-backwards compatible changes. I'm not sure how that's better. Nobody is going to sign on to the first two, so you end up with the exact same situation we have now, but took a more complex approach to get there.
- generationP 7y agoYes, that's what I'm suggesting. I know a big python2 codebase (SageMath) that has been "close" to python3 adaption for several times within the last 5 years but had to call it off every time because something came in the way. By the time the next serious attempt was started, lots of the changes had rotten. Anything that could be done incrementally (python 2.7.5, from-future imports, prints with parentheses) has been done long ago.