4 ms·
I agree with you but when I clicked your footnote source, oh god - the marketing wank. The problem with all of these benchmarks (starting with LAPACK in 1979,
by iheartmemcache 10y ago
I agree with you but when I clicked your footnote source, oh god - the marketing wank. The problem with all of these benchmarks (starting with LAPACK in 1979, up until modern day benchmarks used in the TOP500 or the TPC which models a bank for RDBMS performance) is the synthetic nature of the tests and the unreasonable locality of what they end up testing. One of the bazillions of reasons why the base frequency can drop in those tests Intel used is because your CPU(s) isn't(aren't) context switching or having to do things a normal MSSQL or Oracle DB will do.
I.e., LAPACK/BLAS benchmarks are just really big linear algebra matrix problems, so obviously your pre-fetch and branch prediction performance will be significantly better since you aren't dealing with interrupts, locatedb, or Windows DCOM events firing off in the background. You have a huge set of matrices with a very predictable set of branches, fetches, and decodes, so obviously your CPU can optimize for that load, you're just paying for it in latency on the back-end (RAM fetches are the new disk swap ;)).
On all those benchmarks, (i.e. your standard LU matrix decomposition which previously was the basis of the LAPACK benchmarks, though things might have changed in the ~10 years since I've really looked at things) isn't CPU-bound anymore, so of course your instruction-per-cycle load on the CPU isn't where you'll be bottlenecking (and hasn't been since "let's avoid floating-point operations and just use static look-ups instead since we don't want the 10x cost of using the FDIVP instruction!"). Your processor can very easily anticipate from where in that sparse-matrix your next data fetch is going to be. It's the cost of that RAM fetch[1] going along that copper trace which is going to be where you're going to bottleneck on any heavy numerical computation.
The power consumption on your CPU might drop a nominal amount which is great for those marketing white papers, but for a numerically heavy load, you're paying just as much (in total power consumption per 4U in the data center, total heat generation/dissipation within the case, and total processing time) on the back-end for those fetches.
[1] https://i.stack.imgur.com/a7jWu.png https://i.stack.imgur.com/a7jWu.png (I normally cite academic references, but this is 'good enough' to convey my point, I hope).
- slizard 10y ago> On all those benchmarks, (i.e. your standard LU matrix decomposition which previously was the basis of the LAPACK benchmarks, though things might have changed in the ~10 years since I've really looked at things) isn't CPU-bound anymore, so of course your instruction-per-cycle load on the CPU isn't where you'll be bottlenecking [...] That's incorrect, DGEMM and most BLAS3 is way above most if not all processor uarch's aritmetic intensity threshold [1]. Broadwell CPUs are at 10 Flops/byte [1] while e.g. DGEMM is 32 Flops/byte [2], so that's definitely FLOP/instruction bound and not memory. > Your processor can very easily anticipate from where in that sparse-matrix your next data fetch is going to be. It's the cost of that RAM fetch[1] going along that copper trace which is going to be where you're going to bottleneck on any heavy numerical computation. You're mixing things up, it seems! LAPACK/BLAS is dense matrix, not sparse, so now you switched topics. Sparse matrix ops are generally >1 Flops/byte (see [2]), so that's indeed memory bound. [1] https://www.karlrupp.net/wp-content/uploads/2013/06/flop-per-byte-dp.png https://www.karlrupp.net/wp-content/uploads/2013/06/flop-per... [2] http://www.siam.org/pdf/news/2090.pdf http://www.siam.org/pdf/news/2090.pdf