2 ms·
Its not only CPU cache, its the whole CPU. Having one function per type results in more instructions in total, cause you end up with more copies of the functio
by maxwell86 5y ago
Its not only CPU cache, its the whole CPU.
Having one function per type results in more instructions in total, cause you end up with more copies of the function, but less instructions when dealing with one type, like in a collection of values of one type only.
Less instructions is linear speed up. Single copy of the code enables inlining, constant propagation for type sizes, which enables vectorization of loops. Vectorization alone can buy you up to 32x on modern hardware if the code is flops limited.
Avoiding pointer indirections also trash the cache less, and the BW of caches is orders of magnitude higher than the bandwidth of ram (multiple TB/s vs lower 100 GB/s).
And caches have lower latencies so CPU threads wait less on memory.
So your speed up ends up: less instructions * more FLOPs * more BW, and that can easily be 10x * 32x * 20 ~= 6000x in the worst pathological case.
A 1000x speed up / slowdown per element due to switching between boxed and unboxed generics is less common than a 100x speed up in Rust, but it does happen.