3 ms·
you're benchmarking is python's speed of array instantiation vs. Ds mathematical speed. Its obvious that D will win. Not true, the D code also has the overhead
by bionsuba 11y ago
you're benchmarking is python's speed of array instantiation vs. Ds mathematical speed. Its obvious that D will win.
Not true, the D code also has the overhead of array initialization. std.array.array is called which allocates the results of the range into an array on the GC, which everyone bemoans as being slow as a dog.
Plus, I don't see why this is an invalid benchmark when this is perfectly normal Numpy code, the kind that you see all the time. That was the reason behind the benchmark, to find some common piece of Numpy code and see how equivalent std.ndslice stacked up. I don't see how it's fair to say "Python is really slow at this one thing, so it's not fair to compare it to D in that area".
- cwyers 11y agoBecause, assuming he's right about what's going on, the relationship between complexity and time isn't linear (so D's performance relative to Python will grow less impressive as complexity increases, and at some point Python with NumPy may even perform better), and your example isn't representative of a task that takes long enough for speed to matter.
- semi-extrinsic 11y agoHe's absolutely right, this is a crappy benchmark. Increase the array sizes by at least a factor of 100 to get anything meaningful. The fact that someone who is "the review manager for std.ndslice's inclusion into the standard library" does not understand how to profile numerical algorithms makes me very skeptical of using D for any numerical project. Plus, the syntax looks god-awfully unintuitive. A main advantage of Python is that you often get "code that looks like what it does". The D "basic example with a benchmark" OTOH looks almost obfuscated. To wit; a Fortran version is more readable, is fewer lines of code(!) and of course kicks D's butt when it comes to speed: program p real, dimension(100,1000) :: data real, dimension(1000) :: means n=1 forall(i=1:100,j=1:1000) data(i,j)=n n=n+1 end forall forall(i=1:1000000,j=1:1000) means(j) = sum(data(:,j))/size(data(:,j)) end forall Disclaimer: I wrote this on my phone, only 95% sure that it will compile and run correctly. Save it in means.f90 and compile with `gfortran -Ofast means.f90`. This calculates the means 1 mill. times (the loop over i in the second forall); I bet you it will be an order of magnitude faster than D per means calculation if you time it with plain time (the *nix command).
- bionsuba 11y ago"Increase the array sizes by at least a factor of 100 to get anything meaningful." Ok, let's do that and see what happens: python -m timeit -s 'import numpy; data = numpy.arange(10000000).reshape((1000, 10000))' 'means = numpy.mean(data, axis=0)' D code import std.range : iota; import std.array : array; import std.algorithm; import std.experimental.ndslice; import std.datetime; import std.conv : to; import std.stdio; enum testCount = 10_000; void f0() { auto means = 10_000_000.iota .sliced(1000, 10000) .transposed .map!(r => sum(r) / r.length) .array; } void main() { auto r = benchmark!(f0)(testCount); auto f0Result = to!Duration(r[0] / testCount); f0Result.writeln; } Results Python: 14.1 msec D: 39 μs D is 361.5x faster
- zardeh 11y agoIs D or the compiler doing something sneaky and inlining something, because I'd expect doing 100x more things to take 100x more time for a linear algorithm,but it takes you on the order of 10x more time. Does the original function run in <1ns for you? Edit: Because as I showed above, Pypy managed to JIT inline the entire benchmark, so I wouldn't be surprised at all if a clever compiler managed to do something sneaky.
- semi-extrinsic 11y agoOk. So three points remain: * how do I know that the calculations aren't actually optimized away in all that code I can't understand? E.g. what happens if you make f0() return the means array instead of being a void function? * Given the same number of lines (or characters), are you sure you can't write equally fast Python code? * I tested the Fortran version on a slightly slower computer (got 14.7 msec on the Python version). Fortran runs at 0.7 μs, i.e. > 55x faster than D, with less code that's more readable to boot.
- bionsuba 11y ago"how do I know that the calculations aren't actually optimized away in all that code I can't understand? E.g. what happens if you make f0() return the means array instead of being a void function?" If you can't understand the code, how did you know it was a void function. Please stop with the hyperbole, it's not adding anything to the discussion. Updated code: import std.range : iota; import std.array : array; import std.algorithm; import std.experimental.ndslice; import std.datetime; import std.conv : to; import std.stdio; enum testCount = 10_000; auto f0() { auto means = 10_000_000.iota .sliced(1000, 10000) .transposed .map!(r => sum(r) / r.length) .array; return means; } void main() { auto r = benchmark!(f0)(testCount); auto f0Result = to!Duration(r[0] / testCount); f0Result.writeln; } Results: Python: 14.1 msec D: 41 μs D is 343.9x faster "Given the same number of lines (or characters), are you sure you can't write equally fast Python code?" IMO program size is an almost meaningless statistic outside of code golf challenges. LOC is not an indicative measure of code readability, usefulness, or organization. For example, your Fortran code was 11 lines while the D function (with the return) is eight lines. "I tested the Fortran version on a slightly slower computer (got 14.7 msec on the Python version). Fortran runs at 0.7 μs, i.e. > 55x faster than D, with less code that's more readable to boot." Just goes to show why Fortran is still used in a lot of scientific areas. But two things: 1. Fortran is a much simpler language than D or Python and is much harder to do multipurpose work in it (so I'm told from Fortran programmers, I don't know Fortran myself). So when your program needs to do anything other than number crunching, it's normally done in a separate language. Using D you can have everything in one code base. 2. This article was about Numpy and std.ndslice because those are two areas that I know about and Numpy is a very popular library. Bringing up Fortran's speed here is like commenting on how much faster C++ is in a thread about Ruby. Also, readability is a subjective idea; I believe the D code is more readable than the Fortran you wrote. Different strokes.
- zardeh 11y agoIts that microbenchmarks are useless. If it was worth my time, I could come back and use any of the available python tools (cython, numba, etc.) to compile that function to native code from python syntax and achieve native speed. Or I could write a D implementation of {{some relatively complex matrix algorithm}} and compare it to numpy, showing it was slower. You could then respond by saying that I wasn't using D optimally, and I'd agree.