3 ms·
> On all those benchmarks, (i.e. your standard LU matrix decomposition which previously was the basis of the LAPACK benchmarks, though things might have changed
by slizard 10y ago
> On all those benchmarks, (i.e. your standard LU matrix decomposition which previously was the basis of the LAPACK benchmarks, though things might have changed in the ~10 years since I've really looked at things) isn't CPU-bound anymore, so of course your instruction-per-cycle load on the CPU isn't where you'll be bottlenecking [...]
That's incorrect, DGEMM and most BLAS3 is way above most if not all processor uarch's aritmetic intensity threshold [1]. Broadwell CPUs are at 10 Flops/byte [1] while e.g. DGEMM is 32 Flops/byte [2], so that's definitely FLOP/instruction bound and not memory.
> Your processor can very easily anticipate from where in that sparse-matrix your next data fetch is going to be. It's the cost of that RAM fetch[1] going along that copper trace which is going to be where you're going to bottleneck on any heavy numerical computation.
You're mixing things up, it seems! LAPACK/BLAS is dense matrix, not sparse, so now you switched topics. Sparse matrix ops are generally >1 Flops/byte (see [2]), so that's indeed memory bound.
[1] https://www.karlrupp.net/wp-content/uploads/2013/06/flop-per-byte-dp.png https://www.karlrupp.net/wp-content/uploads/2013/06/flop-per...
[2] http://www.siam.org/pdf/news/2090.pdf http://www.siam.org/pdf/news/2090.pdf