3 ms·
That number my have be true pre 2000, but with the move to parallel architectures it is grossly wrong. Optimization is not about saving a few cycles here and th
by 0x07c0 11y ago
That number my have be true pre 2000, but with the move to parallel architectures it is grossly wrong. Optimization is not about saving a few cycles here and there. It is about using the whole machine, multiple chips, all with multiple cores. This cores are vector cores (AVX/SSE, altivec, etc). The chips are on different NUMA domains, the cores has pipelines, caches, etc, that has to be exploited... Compilers really don't do to much about this. This is now the job of the programmer. Speedup for codes that can exploit this, at least 20-30x. This adds up to loot of computer time... (add multiple accelerators (GPGPUs, MIC), and you can have 200-300x speedup on one node)