6 ms·
Because tuples cannot change over time, they are (slightly) faster to create compared to lists. Like you mentioned, they are being created just to be iterated o
by dosisod 4y ago
Because tuples cannot change over time, they are (slightly) faster to create compared to lists. Like you mentioned, they are being created just to be iterated over. It is a small enough of a performance improvement that it more of a style choice then anything.
Bonus fact: You can use set iteration in for loops as well! This has the added benefit of sorting the values as well:
>>> for x in {2, 1, 3}:
... print(x)
1
2
3
You didn't ask for that, but I felt like sharing it, to there you go
- falcor84 4y agoHere's another tip of syntactic trivia - you can leave the brackets out entirely to (arguably) make it clearer that your don't care about the data structure, and it will implicitly use a tuple: >>> for x in 1, 2, 3: ...
- 5d8767c68926 4y ago>... This has the added benefit of sorting the values as well I don't think you can rely on this? Happy to be proven wrong, but set does not guarantee iteration order. There may be a distinction if all elements are available at set construction, but that seems like a fiddly rule I would rather avoid.
- leetrout 4y agoCareful here- sets are unordered and there is no guarantee when they are or are not going to be ordered as you are demonstrating.
- minhazm 4y agoSets are unordered. It doesn't even make logical sense when you think about it, it would make set insertion O(logn) which would absolutely wreck the performance. Your example just happens to accidentally work. Here's an example where it doesn't: >>> [i for i in {9,40,49}] [40, 9, 49]
- dosisod 4y agoYou're right! I guess you can't rely on sets being iterated over in order (not that you ever should rely on that). Moreover I was trying to show that you can use lists, tuples, and sets as temporary iterable objects, though but no-one (that I have seen) uses sets in that way.
- yunohn 4y agoSets are more useful when you want to dedupe unknown data. Defining a set with constants to use immediately in a for loop is not very helpful compared to a tuple or a list. Also, you can make anything iterable in python if you implement __iter__.
- BeetleB 4y agoSet insertion is O(1). This isn't C++.
- minhazm 4y agoI didn't claim it wasn't O(1). I just said that for sets to automatically be sorted like the parent comment claimed, it would have to be O(logn) insertion since that's the fastest you can insert into a sorted data structure. Obviously sets are not sorted, and thus are able to maintain O(1) insertion.
- BeetleB 4y agoMy bad. Didn't read the context. You're right - what that comment is suggesting is insane.
- cdrini 4y agoFun fact: not only is the iteration order of sets in python unordered, it's also undeterministic! A set of the same values, iterated in two python program, could iterate in a different order. I learned this when a unit test that iterated over a set started failing sporadically :P super confusing bug!
- olejorgenb 4y agoI think this is just for sets whos items hash derive from `hash(some_string)` (eg.: sets of strings) This is because the string hash function is randomized on process startup to help prevent DOS attacks: https://stackoverflow.com/q/30585108/1517969 https://stackoverflow.com/q/30585108/1517969
- evil-olive 4y agothat only works for small numbers. I suspect it's a side effect of the small integer cache CPython maintains [0] >>> for x in {2, 1, 3}: print(x) 1 2 3 >>> for x in {20, 10, 30}: print(x) 10 20 30 >>> for x in {200, 100, 300}: print(x) 200 100 300 0: https://stackoverflow.com/questions/15171695/whats-with-the-integer-cache-maintained-by-the-interpreter https://stackoverflow.com/questions/15171695/whats-with-the-...
- kevincox 4y agoMy guess is that it is because for int `x == hash(x)`. This means that if your integers are smaller than the number of buckets in the hashtable they will fall into buckets in order and the iteration is likely just over the buckets. For example you can see that by adding a large power of two to the numbers you can see that they are sorted by the remainder (as it seems that the default number of buckets is a power of two): > for x in {2048 + 1, 256 + 3, 1024 + 2, 512 + 4}: print(x) 2049 1026 259 516
- cmcconomy 4y agoright, python caches ints from -5 to 256 as singletons, so id(x) for those numbers is constant. whereas each time you run id() on a number outside that range you will get a different id as a new object is instantiated
- kevincox 4y agoI don't think this is related. `set()` uses `hash(x)` it doesn't care what `id(x)` is. The example I provided works equally well with small numbers and large numbers.
- _dain_ 4y ago> Because tuples cannot change over time, they are (slightly) faster to create compared to lists. Like you mentioned, they are being created just to be iterated over. It is a small enough of a performance improvement that it more of a style choice then anything. But that's just wrong. Tuples are not just immutable lists. They're for heterogeneous collections and are meant to be destructured or otherwise have their elements acted on individually. Lists are for homogeneous collections and they're meant to be looped over. Immutability is neither here nor there. The performance benefit is negligible and irrelevant. If you don't believe me, look at how type hints work: an arbitrary-length list of integers is typed like `list[int]`, but you have to write an arbitrary-length tuple of integers like `tuple[int, ...]`. It looks different -- why? Because the way you're meant to use tuples is like `tuple[int, int, str, datetime]`, where there's a fixed number of elements and each one has a distinct meaning, where it usually doesn't make sense to loop over and treat them the same. It's also why there's list comprehension but no tuple comprehension. The suggestion in the example is just plain incorrect.
- vharuck 4y ago>But that's just wrong. Tuples are not just immutable lists. They're for heterogeneous collections and are meant to be destructured or otherwise have their elements acted on individually. Lists are for homogeneous collections and they're meant to be looped over. Immutability is neither here nor there. The performance benefit is negligible and irrelevant. I've never read this opinion before, so I don't think it was intended from the get-go or became a commonly adopted standard. I can see some sense to it, though. But lists aren't just for looping. They often get looped over, but that's true for any collection, even dictionaries. >It's also why there's list comprehension but no tuple comprehension. That's because (x for x in y) was chosen for making generators. Just stick tuple in front and you've got a tuple comprehension.
- _dain_ 4y agoIt's in the official documentation: https://docs.python.org/3/library/stdtypes.html https://docs.python.org/3/library/stdtypes.html >Tuples are immutable sequences, typically used to store collections of heterogeneous data (such as the 2-tuples produced by the enumerate() built-in). Tuples are also used for cases where an immutable sequence of homogeneous data is needed (such as allowing storage in a set or dict instance). GvR email from 2003: https://mail.python.org/pipermail/python-dev/2003-March/033964.html https://mail.python.org/pipermail/python-dev/2003-March/0339... >Tuples are for heterogeneous data, list are for homogeneous data. >Tuples are not read-only lists. Immutable sequence is a use-case, but it's a secondary use-case. The primary use-case is for record-type data. >But lists aren't just for looping. They often get looped over, but that's true for any collection, even dictionaries. My point was more that tuples are not primarily suited for looping over, not that lists exclusively are. >That's because (x for x in y) was chosen for making generators. Just stick tuple in front and you've got a tuple comprehension. But list comprehensions were in the language for several years before generators came around. There was a time when [x for x in foo] worked but (x for x in foo) was a syntax error. Why didn't they make it a tuple at the time? Because that's not what tuples are for. If you look at everywhere tuples show up in the standard library, it's almost always in some context where looping over the collection is not a primary consideration. The clearest example of this dichotomy is in the string methods: the str.split method gives you a list of strings, because the result can be of arbitrary length and you're likely going to loop over it. But str.partition gives a tuple because there are always exactly 3 elements in the result and you're meant to destructure it.
- driggs 4y agoYou are confusing the two different uses of `in`. There's no reason to use a set for a temporary iteration container; but using a set is ideal and - in theory, for large n or expensive equality check - more performant as a temporary containment container: if x in {1, 2, 3}: print(x)