5 ms·
I made a few more interesting (to me) measurements. As always, you have to measure your performance with your actual input data to see what's "best". Test 1:
by davidp 13y ago
I made a few more interesting (to me) measurements. As always, you have to measure your performance with your actual input data to see what's "best".
Test 1: Boring, small array of integers
In [28]: arr = range(0, 300)
In [29]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
10000 loops, best of 3: 27.2 us per loop
In [30]: %timeit numpy.reshape(arr, (-1, 3))
10000 loops, best of 3: 45.2 us per loop
In [31]: %timeit zip(*([iter(arr)]*3))
100000 loops, best of 3: 6.25 us per loop
This roughly matches the article's timing ratios, so far so good.
Test 2: Use numpy's random number generation to get a small array of floats
In [32]: arr = numpy.random.ranf(300)
In [33]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
10000 loops, best of 3: 54 us per loop
In [34]: %timeit numpy.reshape(arr, (-1, 3))
1000000 loops, best of 3: 1.06 us per loop
In [35]: %timeit zip(*([iter(arr)]*3))
10000 loops, best of 3: 39.7 us per loop
numpy is two orders of magnitude faster here; it's evidently using a highly optimized internal codepath for random sequence generation, which I'd guess is a common thing to do in numeric analysis. I assume it's using a generator, so there's no actual array being created, blowing up the CPU cache lines etc.
Test 3: Verify that analysis by interfering with numpy
In [36]: arr = [x for x in numpy.random.ranf(300)]
In [37]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
10000 loops, best of 3: 26.2 us per loop
In [38]: %timeit numpy.reshape(arr, (-1, 3))
10000 loops, best of 3: 48.5 us per loop
In [39]: %timeit zip(*([iter(arr)]*3))
100000 loops, best of 3: 6.55 us per loop
Yep.
Test 4: Larger data set, no interference
In [40]: arr = numpy.random.ranf(3000000)
In [41]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
1 loops, best of 3: 624 ms per loop
In [42]: %timeit numpy.reshape(arr, (-1, 3))
1000000 loops, best of 3: 1.06 us per loop
In [43]: %timeit zip(*([iter(arr)]*3))
1 loops, best of 3: 335 ms per loop
The numpy time doesn't change at all from test 2 despite the larger size, but the others suffer. Again, I suspect numpy is being intelligent here; my guess is that it doesn't actually apply the function and generate the real output, it just wraps the random generator in another one.
Test 5: Larger data set, interfering with numpy
In [44]: arr = [x for x in numpy.random.ranf(3000000)]
In [45]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
1 loops, best of 3: 321 ms per loop
In [46]: %timeit numpy.reshape(arr, (-1, 3))
1 loops, best of 3: 354 ms per loop
In [47]: %timeit zip(*([iter(arr)]*3))
10 loops, best of 3: 83.6 ms per loop
There we go; we're back to roughly the original timing ratios.
So, surprise! You always have to measure. Measure, measure measure. My bias is to write code first for legibility and modifiability, and then optimize hot spots if needed (and add comments, please, when you do so).
Without doing deeper analysis I'd say one moral of the Python story is, this shows the potential power of generators. But in real-world data sets this isn't always ideal -- is it faster to load up the whole data set in memory and blast through it, or load it from disk on demand with a generator? In really high performance scenarios, is it faster to preprocess the data to fit into the CPU's cache lines? You can't tell without measuring, and you have to measure in the environment you're deploying to, since the answer may be different on a machine with 1GB RAM vs. one with 128GB RAM, or 32KB L1 cache vs. 8KB.
- loarake 13y agoThe numpy example becomes fast when you use numpy arrays. Try %timeit numpy.array(arr); numpy.reshape(arr, (-1, 3)); and then just %timeit numpy.array(arr), you'll see that the reshape takes no time at all. Type conversion from python list to numpy array is what kills the performance.
- pudquick 13y agoA point of clarification here - numpy's reshape operation stays fast as long as the array is a numpy array. Which is exactly what the parent comment was all about - the author figured that the reason numpy was significantly faster was because it was accessing / working with the data in a different fashion. So, in order to test that theory, he converted the numpy.array into a normal python array before he proceeded to do any timed operations with zip vs. numpy.reshape, etc. This is a more realistic playing field if you're considering data that was created outside of the numpy environment. At some point, if you're going to work with numpy.reshape, it will need to be type converted / "imported" into numpy data types. For the purposes of this test, it's much more "fair" to include both the time numpy spent on splitting the array as well as that conversion time. The reshape process in numpy had essentially O(1) time with native data types indicating that it had done some behind the scenes work that allowed for such speed. The parent example is much more realistic in capturing the time of the behind the scenes work by forcing each method to start from the same exact same data objects.
- loarake 13y agoMy reply was in response to the statement "numpy is two orders of magnitude faster here; it's evidently using a highly optimized internal codepath for random sequence generation", which is false, it's not because of highly optimized internal codepaths for random sequence generation, it's because the code produced a numpy array (or didn't have to do type conversion). But I agree, when using numpy to produce a timing comparison, it would be fair to start with a numpy array, or to show the time involved in the creation of the array.
- 13y ago