3 ms·
I worked in HPC for over a decade. Most of the codes are not hyper-efficient on Intel - particularly not current-generation Intels that need efficient use of ve
by moconnor 8y ago
I worked in HPC for over a decade. Most of the codes are not hyper-efficient on Intel - particularly not current-generation Intels that need efficient use of very wide vectors to reach peak performance.
Many codes are memory-bound, and the TX2 has excellent bandwidth. This shows up in real-world simulation codes such as OpenFOAM.
- auvi 8y agoVery interesting as I used to do OpenFOAM parallel runs years ago on x86. Do you have any links to OpenFOAM benchmarks on ARM64?
- Ar-Curunir 8y agoThe conclusion of the linked article has a reference to a relevant paper
- gnufx 8y agoBut, as usual, you don't know what the profile of the calculations were, in particular because you don't know the mode it's operating in/data it's operating on. ("OpenFOAM" is actually many different programs.) That said, the indications are that the performance is decent for HPC. What I haven't seen is a comparison with Ryzen (or POWER9). There's also a lack of data even on the SIMD hardware in ThunderX2 and POWER9, at least that I've been able to find.
- gnufx 8y agoIndeed -- memory- or communication-bound. Somewhere there's a Dell "roofline" plot indicating IPC in some codes running in an unspecified way. See also John McCalpin's writings. Things doing 3-D FFTs (common in materials science) tend to spend a lot of time in an MPI collective at sufficient scale. I spent half an hour profiling and then changing an MPI parameter for 30% improvement in one case at not very large scale.