6 ms·
Keyme, you and I are the few that have serious concern with the design of Python 3. I started to embrace it in a big way 6 months ago when most libraries I use
by tungwaiyip 13y ago
Keyme, you and I are the few that have serious concern with the design of Python 3. I started to embrace it in a big way 6 months ago when most libraries I use are available in Python. I wish to say the wait is over and we should all move to Python 3 then everything will be great. Instead I find no compelling advantage. Maybe there will be when I start to use unicode string more. Instead I'm really annoyed by the default iterator and the binary string handling. I am afraid it is not a change for the good.
I come from the Java world when people take a lot of care to implement things as streams. It was initially shocking to see Python read an entire file into memory, turn it into list or other data structure with no regard to memeory usage. Then I have learned this work perfectly well when you have a small input, a few MB or so is a piece of cake for modern computer. It takes all the hassle out of setting up streams in Java. You optimize when you need to. But for 90% of stuff, a materialized list works perfectly well.
Now Python become more like Java in this respect. I can't do exploratory programming easily without adding list(). Many times I run into problem when I am building complex data structure like list of list, and end up getting a list of iterator. It takes the conciseness out of Python when I am forced to deal with iterator and to materialize the data.
The other big problem is the binary string. Binary string handling is one of the great feaute of Python. It it so much more friendly to manipulate binary data in Python compare to C or Java. In Python 3, it is pretty much broken. It would be an easy transition I only need to add a 'b' prefix to specify it as binary string literal. But in fact, the operation on binary string is so different from regular string that it is just broken.
In [38]: list('abc')
Out[38]: ['a', 'b', 'c']
In [37]: list(b'abc') # string become numbers??
Out[37]: [97, 98, 99]
In [43]: ''.join('abc')
Out[43]: 'abc'
In [44]: ''.join(b'abc') # broken, no easy way to join them back into string
---------------------------------------------------------------------------
TypeError Traceback (most recent call last)
<ipython-input-44-fcdbf85649d1> in <module>()
----> 1 ''.join(b'abc')
TypeError: sequence item 0: expected str instance, int found
- maxerickson 13y ago>>> bytes(list(b'abc')) b'abc' >>> That is, the way to turn a list of ints into a byte string is to pass it to the bytes object. (This narrowly addresses that concern, I'd readily concede that the new API is going to have situations where it is worse)
- tungwaiyip 13y agoThank you. This is good to know. I was rather frustrated to find binary data handling being changed with no easy translation in Python 3. Here is another annoyance: In [207]: 'abc'[0] + 'def' Out[207]: 'adef' In [208]: b'abc'[0] + b'def' --------------------------------------------------------------------------- TypeError Traceback (most recent call last) <ipython-input-208-458c625ec231> in <module>() ----> 1 b'abc'[0] + b'def' TypeError: unsupported operand type(s) for +: 'int' and 'bytes'
- maxerickson 13y agoI don't have enough experience with either version to debate the merits of the choice, but the way forward with python 3 is to think of bytes objects as more like special lists of ints, where if you want a slice (instead of a single element) you have to ask for it: >>> [1,2,3][0] 1 >>> [1,2,3][0:1] [1] >>> b'abc'[0] 97 >>> b'abc'[0:1] b'a' >>> So the construction you want is just: >>> b'abc'[0:1]+b'def' b'adef' Which is obviously worse if you are doing it a bunch of times, but it is at least coherent with the view that bytes are just collections of ints (and there are situations where indexing operations returning an int is going to be more useful).
- tungwaiyip 13y agoIn Java, String and Char are two separate types. In Python, there is no separate char type. It is simply a string of length of 1. I do not have great theory to show which design is better either. I can only say the Python design work great for me in the past (for both text and binary string), and I suspect it is the more user friendly design of the two. So in Python 3 the design of binary string is changed. Unlike the old string, bytes and binary string of length 1 are not the same. Working codes are broken, practice have to be changed, often it involves more complicated code (like [0] becomes [0:1]). All these happens with no apparent benefit other than it is more "coherent" in the eye of some people. This is the frustration I see after using Python 3 for some time.
- keyme 13y agoYes! Thank you. All the other commenters here that are explaining things like using a list() in order to print out an iterator are missing the point entirely. The issue is "discomfort". Of course you can write code that makes everything work again. This isn't the issue. It's just not "comfortable". This is a major step backwards in a language that is used 50% of the time in an interactive shell (well, at least for some of us).
- agentultra 13y agoThe converse problem is having to write iterator versions of map, filter, and other eagerly-evaluated builtins. You can't just write: >>> t = iter(map(lambda x: x * x, xs)) Because the map() call is eagerly evaluated. It's much easier to exhaust an iterator in the list constructor and leads to a consistent iteration API throughout the language. If that makes your life hard then I feel sorry for you, son. I've got 99 problems but a list constructor ain't one. The Python 3 bytes object is not intended to be the same as the Python 2 str object. They're completely separate concepts. Any comparison is moot. Think of the bytes object as a dynamic char[] and you'll be less inclined to confusion and anger: >>> list(b'abc') [97, 98, 99] That's not a list of numbers... that's a list of bytes! >>> "".join(map(lambda byte: chr(byte), b'abc')) 'abc' And you get a string!
- itsadok 13y ago> The converse problem is having to write iterator versions of map, filter, and other eagerly-evaluated builtins Well, in Python 2 you just use imap instead of map. That way you have both options, and you can be explicit rather than implicit. > That's not a list of numbers... that's a list of bytes! The point being made here is not that some things are not possible in Python 3, but rather than things that are natural in Python 2 are ugly in 3. I believe you're proving the point here. The idea that b'a'[0] == 97 in such a fundamental way that I might get one when I expected the other may be fine in C, but I hold Python to a higher standard.
- dded 13y ago> >>> list(b'abc') > [97, 98, 99] >That's not a list of numbers... that's a list of bytes! No, it's a list of numbers: >>> type(list(b'abc')[0]) <class 'int'> I think the GP mis-typed his last example. First, he showed that ''.join('abc') takes a string, busts it up, then concatenates it back to a string. Then, with ''.join(b'abc'), he appears to want to bust up a byte string and concatenate it back to a text string. But I suspect he meant to type this: >>> b''.join(b'abc') That is, bust up a byte string and concatenate back to what you start with: a byte string. But that doesn't work, when you bust up a byte string you get a list of ints; and you cannot concatenate them back to a byte string (at least not elegantly).