4 ms·
Thanks for sharing the benchmarks. I made the arrays huge to get roughly the same ball park of times you reported. I also preallocated the memory in that bench
by celrod 6y ago
Thanks for sharing the benchmarks.
I made the arrays huge to get roughly the same ball park of times you reported.
I also preallocated the memory in that benchmark to avoid unnecessary allocations (the observed allocations are because multithreading in Julia allocates memory). Is this something as easy to do in Python?
I made a few tweaks to your benchmark code. I get with Julia:
Element-wise sum of two 100x100 matrices (int), (1000 loops)
1.698 ms (2 allocations: 78.20 KiB)
Element-wise multiplication of two 100x100 matrices (float64), (1000 loops)
1.786 ms (2 allocations: 78.20 KiB)
Dot (scalar) product of two 300000 arrays (float64), (1000 loops)
4.801 ms (0 allocations: 0 bytes)
Matrix product of 500x600 and 600x500 matrices (float64)
229.475 μs (0 allocations: 0 bytes)
L2 norm of 500x600 matrix (float64), (1000 loops)
6.566 ms (0 allocations: 0 bytes)
Sort of 500x600 matrix (float64)
523.243 μs (595 allocations: 2.32 MiB)
Python:
| Element-wise sum of two 100x100 matrices (int), (1000 loops) | 0.0042680892984208185 |
| Element-wise multiplication of two 100x100 matrices (float64), (1000 loops) | 0.0027007726486772297 |
| Dot (scalar) product of two 300000 arrays (float64), (1000 loops) | 0.0088505959516624 |
| Matrix product of 500x600 and 600x500 matrices (float64) | 0.00107754109922098 |
| L2 norm of 500x600 matrix (float64), (1000 loops) | 0.010691180499998154 |
| Sort of 500x600 matrix (float64) | 0.008244920199649642 |
That is, Julia-speedup factors of:
1. Small int matrix element-wise sum: 2.5x
2. Small float matrix element-wise product: 1.5x
3. Dot product (scalar): 1.84x
4. Matrix multiplication: 4.7x
5. L2 norm: 1.6x
6. Sorting: 15.8x
The changes were:
1. Memory allocation was slow, but Julia makes in-place operations easy, so I made 1, 2, and 4 in-place. Although I still allocate the result within the test function for `1.` and `2.`.
2. MKL is much faster than OpenBLAS, so I made `BLAS.vendor()` return `:mkl`
3. I sorted in a threaded loop and transposed the destination array. Transposing the destination array was because Julia is column major, so sorting rows is a benchmark that will naturally favor the row-major language.
But I left all the inputs the same.
My primary point with Julia here is that it's easy to optimize it and get more performance out of it if or whenever you need it.