4 ms·
I just noticed your slotted object uses class variables. That is completely wrong - so your analysis is probably wrong. In Python, I think it's important to do
by aaronchall 9y ago
I just noticed your slotted object uses class variables. That is completely wrong - so your analysis is probably wrong.
In Python, I think it's important to do what's directly semantically correct.
However, we can compare apples and oranges on what they have in common, so here's a bit of my own analysis:
>>> from timeit import repeat
>>> from collections import namedtuple
>>> class Slotted: __slots__ = 'a', 'b', 'c', 'd', 'e'
...
>>> NT = namedtuple('NT', 'a b c d e')
>>>
>>> nt = NT(1,2,3,4,5)
>>> s = Slotted()
>>> s.a, s.b, s.c, s.d, s.e = 1,2,3,4,5
>>> min(repeat(lambda: (s.a, s.b, s.c, s.d, s.e)))
0.2850431999977445
>>> min(repeat(lambda: (nt.a, nt.b, nt.c, nt.d, nt.e)))
0.418955763001577
I'm showing the namedtuple to be a little slower, but I think the thing to remember is to do the semantically correct thing first. If you've met your technical requirements, you're done.
Only if you're too slow, then you look for ways to optimize for bottlenecks.
- gshulegaard 9y agoI was just copying the benchmark I was quoting as a sanity check since I was just trying to validate references. I also feel compelled to point out that __slots__ changes the way objects are initialized [1]: > By default, instances of both old and new-style classes have a dictionary for attribute storage. This wastes space for objects having very few instance variables. The space consumption can become acute when creating large numbers of instances. > > The default can be overridden by defining __slots__ in a new-style class definition. The __slots__ declaration takes a sequence of instance variables and reserves just enough space in each instance to hold a value for each variable. Space is saved because __dict__ is not created for each instance. Which means that, in actuality, simply defining __slots__ changes the nature of the object. Quick trial demonstrates this: >>> class mySlottedObj(object): ... __slots__ = ('a', 'b') ... c = 1 ... >>> x = mySlottedObj() >>> x <mySlottedObj object at 0x10e6a3510> >>> x.c 1 >>> x.__dict__ Traceback (most recent call last): File "<input>", line 1, in <module> AttributeError: 'mySlottedObj' object has no attribute '__dict__' >>> x.__weakref__ Traceback (most recent call last): File "<input>", line 1, in <module> AttributeError: 'mySlottedObj' object has no attribute '__weakref__' >>> x.a = 1 >>> x.c = 2 Traceback (most recent call last): File "<input>", line 1, in <module> AttributeError: 'mySlottedObj' object attribute 'c' is read-only The benefit of namedtuples, in the context of Python is reducing the object footprint to being no greater than a tuple. The performance hit, which I am trying to demonstrate, is in lookup and is shown in the source (linked above). _field_template = '''\ {name} = _property(_itemgetter({index:d}), doc='Alias for field number {index:d}') ''' As I understand it this means when you access using __getitem__, a named tuple first maps to a property that then calls a __getitem__ using the aliased index on the base `tuple` object that `namedtuple` is storing. But at any rate, the benchmark script was not mine and refining the test to be as close to the same code path is welcome. And I wholeheartedly agree: > Only if you're too slow, then you look for ways to optimize for bottlenecks. I was merely responding to the original assertion: > This provides the same performance benefits, and is very similar in a lot of ways to a Scala case class -- aside from removing the boilerplate of stating each attribute twice in __init__ and again in __slots__, And I think at this point it is quite clear that namedtuple has a bit more caveats than just "removing the boilerplate" of __slots__.