3 ms·
I have to admit, I was skeptical but my complier did pick SIMD for the second and not the first and it did make a huge difference. % cat t.c #include <sys/
by esmi 8y ago
I have to admit, I was skeptical but my complier did pick SIMD for the second and not the first and it did make a huge difference.
% cat t.c
#include <sys/time.h>
#include <sys/resource.h>
#include <stdio.h>
double get_time()
{
struct timeval t;
struct timezone tzp;
gettimeofday(&t, &tzp);
return t.tv_sec + t.tv_usec*1e-6;
}
int blah[SIZE][SIZE];
int whatever() { static int i=0; return ++i; }
int main()
{
double t0 = get_time();
for(int i=0; i<SIZE; i++){
for(int j=0; j<SIZE; j++){
blah[j][i] = whatever(); // Column oriented is very slow.
}
}
double t1 = get_time();
for(int j=0; j<SIZE; j++){
for(int i=0; i<SIZE; i++){ // Very fast: We move row-wise now. With luck, this is SIMD-vectorized by your compiler.
blah[j][i] = whatever();
}
}
double t2 = get_time();
printf("SIZE=%5d dt1=%4.3e dt2=%4.3e\n",SIZE,t1-t0,t2-t1);
return 0;
}
% for (( i=1; i<100000 ; i*=10 )) do gcc -DSIZE=$i -O t.c; ./a.out; done
SIZE= 1 dt1=4.179e-04 dt2=0.000e+00
SIZE= 10 dt1=5.190e-04 dt2=0.000e+00
SIZE= 100 dt1=5.062e-04 dt2=9.537e-07
SIZE= 1000 dt1=4.014e-03 dt2=1.490e-04
SIZE=10000 dt1=1.347e+00 dt2=4.349e-02
2.8G Core i7 1600MHz DDR3
- WallWextra 8y agon.b. there would be a difference even without vectorization
- gdy 8y agoNow you could try to time it with -fno-tree-vectorize or whatever gcc uses nowadays. You might be surprised again :)
- dragontamer 8y agoYup yup. SIMD is a CPU-level optimization, but the problem (at size 10,000 x 10,000) is DDR4 memory-limited. SIMD only really makes a difference at the L1 or L2 cache levels. Its a fancy micro-optimization that compilers do and can improve code speed in those cases... But the "big" change, going from column-wise traversal into row-wise traversal, is the huge memory optimization that programmers should know about. It just so happens that SIMD-optimizations are also easier for compilers to figure out on row-wise traversal, so you get SIMD-optimization "for free" in many cases.