24 ms·
In C99, you can use the "restrict" keyword to tell the compiler an array is not aliased, and even without restrict it's often possible to get the same optimisat
by smarnach 10y ago
In C99, you can use the "restrict" keyword to tell the compiler an array is not aliased, and even without restrict it's often possible to get the same optimisations in C or C++. When I was still doing HPC in academia, all our C++ code was maxing out the memory bandwidth. I think in practice the difference rarely matters these days.
- bmh100 10y agoWere you using C99 during your HPC programming in academia?
- smarnach 10y agoI used C99 for the computing kernel of one standalone project (simulation of gravitational lensing). Apart from this, it was mostly C++.
- bmh100 10y agoWhat made your work so dependent upon memory bandwidth?
- greglindahl 10y agoAs an introduction, you might want to look into John McCalpin's history as a shallow-water oceanographer and his subsequent fixation on the STREAM benchmark. TL;DR: some algorithms aren't friendly to caching. I have no idea if this applies to smarnach, but it definitely applies to other people.
- jabl 10y agoJohn McCalpin did a nice presentation at SC16 giving an overview of memory BW/latency issues in HPC, slides: http://sites.utexas.edu/jdm4372/2016/11/22/sc16-invited-talk-memory-bandwidth-and-system-balance-in-hpc-systems/ http://sites.utexas.edu/jdm4372/2016/11/22/sc16-invited-talk...
- flamedoge 10y agoyes, but restrict is limited in that it can only be used to tell two pointed memories could never alias. Finer grain aliasing can be done using unions and structs, but it's not as flexible as Fortran and the possibility of non-aliasing is easily broken just by allowing address to be taken.
- greglindahl 10y agoInsomuch as I've seen HPC code written by science academics, both when I was one and later when I became an HPC guy, as much as I cringe at some of the Fortran, the C/C++ can be far worse... and I wish that I could retroactively restrict some of my colleagues and customers to sticking with Fortran.
- greglindahl 10y agoWow, look at the downvotes! I've been saying this to people for 20+ years, in conference talks and face-to-face conversations, and I've never had a negative reaction. Happy to learn if folks think I'm wrong about my personal experiences. If you've had a different experience, happy to learn about that too (although I'd be surprised if you'd downvote me for that.) (edit: thanks for the upvotes! and discussion!)
- GFK_of_xmaspast 10y agoPersonally I'd rather see academic c than academic fortran; academic c++ is in the 'oh please no' category.
- greglindahl 10y agoYes, I've seen too many unfortunate template libraries, which never work in combination with a different template library.
- ryescienceguy 10y agoI've used some quite reasonable academic C++ for geometry processing and meshing.
- greglindahl 10y agoNote that I'm not saying that science academic C++ is never great; I'm saying that it's often unfortunate. Is the code you're talking about code that's progressed through several maintainers without becoming sad? That's often where things go wrong, not the initial code, but keeping it going, cleanly, through multiple collaborators / post-docs / grad students. I agree that meshing is definitely a difficult situation; I've never seen a solution (for unusual geometry) that I liked much. Things that can pretend to be rectangular? No problem. Very adaptive meshes? Argh.
- new299 10y agoAs an aside, how to you accurately measure memory bandwidth? I've been looking for a good solution but couldn't find one.
- pandaman 10y agoSome processors expose performance counters but you don't need to accurately measure memory bandwidth to know you are maxing out. Suppose you know that the memory bandwidth is 1Gb/s and you program reads 100Gb of data to produce the result. If the execution takes ~100s then you know you hit the maximum memory bandwidth.
- deleted 10y ago[deleted]
- kazinator 10y agoYou can also quite simply use local variables, and unroll loops manually: for (i ...) { double a0 = a[i]; double a1 = a[i+1]; double b0 = b[i]; double b1 = b[1+1]; c[i] = a0 + b0; c[1+1] = a1 + b1; } The hand-unrolling is predicated on the knowledge that the arrays do not overlap. It is obvious to the compiler that the locals do not alias with the arrays, or with each other. The stores to c[] could alias with a[] or b[], but that hardly matters; we assume that all the array loads and stores take place. We have done ourselves the unrolling that was hindered by the suspicion that the arrays overlap, and could do more of it. When coding in C for tightness, it helps to pretend that locals (at least basic type scalar ones) are registers and references to arrays and structs are memory loads and stores. Then think like an assembly language programmer, somewhat. It's better to use a few more locals than a few more loads and stores. Try to load once, process with locals, store once.