3 ms·
It's worth noting that HPC is a different problem space from say, kernels, drivers, codecs, and browsers, in that HPC is not generally expecting to deal with ad
by Robin_Message 3y ago
It's worth noting that HPC is a different problem space from say, kernels, drivers, codecs, and browsers, in that HPC is not generally expecting to deal with adversairial input, and those other spaces are. So the safety-performance trade-off is truly different between you and many of the commentators who are rightly pointing out C/C++'s atrocious track record in the secure space.
Also, I've done a small amount of scientific HPC, and as I see it "correct" is infinitely more important than "fast". If you look at the amount of incorrect scientific papers that trace back to accidentally and silently corrupted data, I think maybe it might sense to consider using any and all possible tools to avoid corruption, of which Rust is but one example (and you've named others like valgrind).
- bayindirh 3y agoHowever, this doesn't mean that an HPC application has no need to verify its input, fail gracefully if something goes wrong, or have to be absolutely rock solid because it's running for days (or weeks in some cases) while being dead-on the results of the problem evaluated. You're right. Correct is the king in HPC, but it's a king which doesn't exile speed. What we do is to solve a couple of known cases correctly in MATLAB or something similar, then try to hit the same results first (or be better if we can do that), then make it iteratively faster without deviating it from a bit (Same result with 1e-32 precision, IOW). This "comparing with a known truth" process is a great indicator of calculation sanity. Other parts are iteratively tested with Valgrind and custom test suites, from function level to end to end, automatically after each build. End to end tests are very expensive in the Valgrind memory profiler (esp. if you're testing multi-thread code), so we do these weekly generally, but I think the idea is clear.