4 ms·
I tried element-wise sum, and got julia> using LoopVectorization, BenchmarkTools julia> A = rand(10_000, 10_000); B = rand(10_000, 10_000); C = similar(A)
by celrod 6y ago
I tried element-wise sum, and got
julia> using LoopVectorization, BenchmarkTools
julia> A = rand(10_000, 10_000); B = rand(10_000, 10_000); C = similar(A);
julia> @benchmark vmapntt!(+, $C, $A, $B)
BenchmarkTools.Trial:
memory estimate: 13.19 KiB
allocs estimate: 91
--------------
minimum time: 29.103 ms (0.00% GC)
median time: 29.202 ms (0.00% GC)
mean time: 29.329 ms (0.00% GC)
maximum time: 30.658 ms (0.00% GC)
--------------
samples: 171
evals/sample: 1
While with NumPy
>>> import timeit
>>> u = timeit.Timer("A + B", setup='import numpy as np; A = np.random.rand(10_000, 10_000); B = np.random.rand(10_000,10_000)')
>>> u.repeat(10, 1)
[0.17918606100283796, 0.17888473700440954, 0.17893354399711825, 0.1790916720055975, 0.17922663199715316, 0.17935074500564951, 0.17939399600436445, 0.17940548500337172, 0.17933111900492804, 0.17920518800383434]
A 6x advantage for Julia. If I don't use multithreading, Julia slows down to 122.5ms, or about 1.5x faster.
Elementwise multiplication is the same fast.
If I get a little more creative and do
exp(A) + log(B) instead, multithreaded Julia still takes 30ms, while single threaded slows down to 169ms.
Meanwhile, NumPy slowed down to 847ms, making it 30x slower than multithreaded Julia (on my computer) and 5x slower than single threaded.
Shall I keep making programs more complicated?
- tastyminerals2 6y agoNever use big arrays in benchmarks otherwise depending on your CPU cache you will be testing allocation as well. If you need the code I tested Julia on here you go: https://github.com/tastyminerals/mir_benchmarks/blob/master/other_benchmarks/julia_bench.jl https://github.com/tastyminerals/mir_benchmarks/blob/master/...
- celrod 6y agoThanks for sharing the benchmarks. I made the arrays huge to get roughly the same ball park of times you reported. I also preallocated the memory in that benchmark to avoid unnecessary allocations (the observed allocations are because multithreading in Julia allocates memory). Is this something as easy to do in Python? I made a few tweaks to your benchmark code. I get with Julia: Element-wise sum of two 100x100 matrices (int), (1000 loops) 1.698 ms (2 allocations: 78.20 KiB) Element-wise multiplication of two 100x100 matrices (float64), (1000 loops) 1.786 ms (2 allocations: 78.20 KiB) Dot (scalar) product of two 300000 arrays (float64), (1000 loops) 4.801 ms (0 allocations: 0 bytes) Matrix product of 500x600 and 600x500 matrices (float64) 229.475 μs (0 allocations: 0 bytes) L2 norm of 500x600 matrix (float64), (1000 loops) 6.566 ms (0 allocations: 0 bytes) Sort of 500x600 matrix (float64) 523.243 μs (595 allocations: 2.32 MiB) Python: | Element-wise sum of two 100x100 matrices (int), (1000 loops) | 0.0042680892984208185 | | Element-wise multiplication of two 100x100 matrices (float64), (1000 loops) | 0.0027007726486772297 | | Dot (scalar) product of two 300000 arrays (float64), (1000 loops) | 0.0088505959516624 | | Matrix product of 500x600 and 600x500 matrices (float64) | 0.00107754109922098 | | L2 norm of 500x600 matrix (float64), (1000 loops) | 0.010691180499998154 | | Sort of 500x600 matrix (float64) | 0.008244920199649642 | That is, Julia-speedup factors of: 1. Small int matrix element-wise sum: 2.5x 2. Small float matrix element-wise product: 1.5x 3. Dot product (scalar): 1.84x 4. Matrix multiplication: 4.7x 5. L2 norm: 1.6x 6. Sorting: 15.8x The changes were: 1. Memory allocation was slow, but Julia makes in-place operations easy, so I made 1, 2, and 4 in-place. Although I still allocate the result within the test function for `1.` and `2.`. 2. MKL is much faster than OpenBLAS, so I made `BLAS.vendor()` return `:mkl` 3. I sorted in a threaded loop and transposed the destination array. Transposing the destination array was because Julia is column major, so sorting rows is a benchmark that will naturally favor the row-major language. But I left all the inputs the same. My primary point with Julia here is that it's easy to optimize it and get more performance out of it if or whenever you need it.