4 ms·
As I stated in the article, I did not include the array creation in the benchmark in order to be fair to Numpy with its slow initialization times. The only pyth
by bionsuba 11y ago
As I stated in the article, I did not include the array creation in the benchmark in order to be fair to Numpy with its slow initialization times. The only python code that I benchmarked was the numpy.mean line.
- zardeh 11y agoThat still doesn't matter. So much of the work is being done in python that you're benchmarking the overhead of python, not the numpy speed. I mean this is a terrible benchmark. Comparing to pure python, the numpy code is about 4-5x faster on my machine, whereas in reasonable real world benchmarks, numpy is hundreds of times faster than pure python. For reference my benchmark was %timeit [sum(row)/len(row) for row in mat] and it completed in about 900us, given mat was a python 2d array. These results are as valuable as the ones pypy gave me: pypy -mtimeit -s "[sum(row)/len(row) for row in [[range(100000)[i*j] for i in range(100)] for j in range(1000)]]" 1000000000 loops, best of 3: 0.00101 usec per loop Edit: In other words, because the code returns a python array of length 1000, what you're benchmarking is python's speed of array instantiation vs. Ds mathematical speed. Its obvious that D will win. The only fair way to benchmark the mathematics here is if you compare the D implementation of O(n) algorithm that returns a single value, or an O(n^2) algorithm that returns a 1d array or single value etc. Otherwise the python overhead of creating O(n) objects will overshadow the actual mathematical calculations on those O(n) objects.
- bionsuba 11y agoyou're benchmarking is python's speed of array instantiation vs. Ds mathematical speed. Its obvious that D will win. Not true, the D code also has the overhead of array initialization. std.array.array is called which allocates the results of the range into an array on the GC, which everyone bemoans as being slow as a dog. Plus, I don't see why this is an invalid benchmark when this is perfectly normal Numpy code, the kind that you see all the time. That was the reason behind the benchmark, to find some common piece of Numpy code and see how equivalent std.ndslice stacked up. I don't see how it's fair to say "Python is really slow at this one thing, so it's not fair to compare it to D in that area".
- cwyers 11y agoBecause, assuming he's right about what's going on, the relationship between complexity and time isn't linear (so D's performance relative to Python will grow less impressive as complexity increases, and at some point Python with NumPy may even perform better), and your example isn't representative of a task that takes long enough for speed to matter.
- semi-extrinsic 11y agoHe's absolutely right, this is a crappy benchmark. Increase the array sizes by at least a factor of 100 to get anything meaningful. The fact that someone who is "the review manager for std.ndslice's inclusion into the standard library" does not understand how to profile numerical algorithms makes me very skeptical of using D for any numerical project. Plus, the syntax looks god-awfully unintuitive. A main advantage of Python is that you often get "code that looks like what it does". The D "basic example with a benchmark" OTOH looks almost obfuscated. To wit; a Fortran version is more readable, is fewer lines of code(!) and of course kicks D's butt when it comes to speed: program p real, dimension(100,1000) :: data real, dimension(1000) :: means n=1 forall(i=1:100,j=1:1000) data(i,j)=n n=n+1 end forall forall(i=1:1000000,j=1:1000) means(j) = sum(data(:,j))/size(data(:,j)) end forall Disclaimer: I wrote this on my phone, only 95% sure that it will compile and run correctly. Save it in means.f90 and compile with `gfortran -Ofast means.f90`. This calculates the means 1 mill. times (the loop over i in the second forall); I bet you it will be an order of magnitude faster than D per means calculation if you time it with plain time (the *nix command).
- bionsuba 11y ago"Increase the array sizes by at least a factor of 100 to get anything meaningful." Ok, let's do that and see what happens: python -m timeit -s 'import numpy; data = numpy.arange(10000000).reshape((1000, 10000))' 'means = numpy.mean(data, axis=0)' D code import std.range : iota; import std.array : array; import std.algorithm; import std.experimental.ndslice; import std.datetime; import std.conv : to; import std.stdio; enum testCount = 10_000; void f0() { auto means = 10_000_000.iota .sliced(1000, 10000) .transposed .map!(r => sum(r) / r.length) .array; } void main() { auto r = benchmark!(f0)(testCount); auto f0Result = to!Duration(r[0] / testCount); f0Result.writeln; } Results Python: 14.1 msec D: 39 μs D is 361.5x faster
- srean 11y agoYou may have a point that the benchmarks do not rule out the possibility that D code might be asymptotically slower. But benchmarks that are designed to hide the fact that CPython function call overheads and object creation overheads are a godawful abomination are far from a fair benchmark. Those overheads are real and have real effects.
- zardeh 11y agoSure, but the claim he's making is that D is faster than the numpy numerical algorithms. This is patently false. The C or fortan code that is doing the numerical calculation is as fast or faster than the D code he's running. Al of the difference comes from the initialization of O(n) pyobjects to represent the returned values. If he claimed that handwritten D code was faster than equivalent python, I'd say "duh", but he claimed that the numerical libraries are faster. I take issue to the statement that D will be faster than numpy for real world calculations, because the numpy code will run the actual math faster, and the python overhead will still be on the order of nano or milliseconds.