4 ms·
Vectorization speeds up your Python code
- rytill 5y agoGreat article. I especially liked the comparison of float32 to float64. I had no idea it was as large as a 10x improvement in a simple case like the one demonstrated. * Elapsed (secs) * CPU instructions (G) * Cache misses (M) * float64 : 4.7 1.6 30.4 * float32 : 0.4 1.2 16.7
- dannyz 5y agoThis looks quite odd to me. I don't have perf installed on my machine but when I do DATA = np.random.rand(300_000_000).astype(np.float32) I get ~0.24s as a wall time for the normalization calculation on my machine, and DATA = np.random.rand(300_000_000).astype(np.float64) Is giving me ~0.33s
- itamarst 5y agoNote that line needs to be omitted from the time measurement, since it's the same for both scripts and is just overhead, I was just measuring the actual mean + substraction code. I reran a few more times, got a bunch of variability, but did get some runs where f64 is just twice as slow, not 10×. Probably my initial run hit swapping. So I will update the article. That being said, worth noting that this is also quite hardware dependent. Like, I just have 16MB L3 cache, I imagine higher L3 would change ratios a lot. UPDATE: Oops. Original runs were with NumPy 1.18, this run was with NumPy 1.22. So will rerun again with that. UPDATE2: NumPy version doesn't matter. So yeah was probably swapping. UPDATE3: Fixed article should be up in a minute.