3 ms·
My guess/intuition was that walking backwards might defeat any memory pre-caching at various layers (CPU, OS). I wrote this quick-and-dirty code [http://codepa
by jbert 13y ago
My guess/intuition was that walking backwards might defeat any memory pre-caching at various layers (CPU, OS).
I wrote this quick-and-dirty code [http://codepad.org/d9ilykxA http://codepad.org/d9ilykxA] to try and test the effect.
To me, it's a little interesting how it various with optimisation level, but I've not picked apart the assembly to see what's going on.
Under gcc 4.8.1, 64bit ubuntu):
-O0 forwards 4.37s, backwards 4.88s (fowards faster)
-O1 forwards 2.67s, backwards 2.66s (same/backwards faster)
-O2 forwards 2.67s, backwards 2.66s (same/backwards faster)
-Os forwards 3.33s, backwards 3.00s (backwards faster)
and the same under clang:
-O0 forwards 3.35s, backwards 4.11s
-O1 forwards 2.67s, backwards 2.67s
-O2 forwards 0.69s, backwards 0.77s
-Os forwards 2.67s, backwards 3.00s
So...it's not as simple as "forwards or backwards always faster", unless the test code is simple enough to be defeated by the optimiser in some cases which real code wouldn't.
Also - what is going on with clang -O2?
- goldenkey 13y agoI don't believe these tests show much. You need to be doing something realistic inside the loop, and not repeating it using num_iterations, that will probably be optimized out into an unrolled memcpy. Too simplistic to be an accurate test. Modern CPUs will precache memory forwards as well as in reverse.