4 ms·
This is a very nice writeup, and clearly dictates performance benefits to be gained from using underlying hardware properly. Couple of things i noticed. ``` #
by warangal 3y ago
This is a very nice writeup, and clearly dictates performance benefits to be gained from using underlying hardware properly.
Couple of things i noticed.
```
# calculating dot product
a = np.random.randn(1,1536).astype("float32")
b = np.random.randn(1536, 1).astype("float32")
result = np.matmul(a * b) ~1.1 microsecond
```
My system configurations are: Intel(R) Core(TM) i5-8300H CPU @ 2.30GHz 2.30 GHz, (quad core) AVX256 instructions are supported.
problem lies in calculating the magnitude of vectors, which i guess scipy is doing as ``np.sum(np.square(a)``.Any numpy routine like ``matmul`` using Blas library would be very very fast, but all other routines are just simple C for loops. Magnitude calculation by passing numpy buffer to a C extension takes about ``1.4`` microseconds. using SIMD instructions for magnitude calculation is very straightforward and i get a speed up of about 8x.
For critical calculations you are almost always better off by passing numpy buffer to C code and fusing operations there, and optionally speed up code using SIMD instructions based on hardware you want to target.
It should be ``scipy.spatial.distance.cosine``, not ``scipy.distance.spatial.cosine`` in your blog !