4 ms·
A reason for the slow adoption of accelerators is that these code bases are pretty old and mostly Fortran based. I don't think we are going to see a shift to a
by __mp 10y ago
A reason for the slow adoption of accelerators is that these code bases are pretty old and mostly Fortran based.
I don't think we are going to see a shift to accelerator-based weather models any time soon. The code bases are pretty old and most scientists working on these models don't have the experience or do not want to work in a programming methodology that works well with accelerators (at least that's my experience). Also in a lot of cases people just like to work with an already existing code because they are just doing a PhD and don't want to translate everything to a new language.
The COSMO model which I'm working on is 18 years old. We ported the compute intensive part of the model to C++ using a stencil library and integrated OpenACC pragmas (think OpenMP, but for GPUs) in the rest of the code base.
I'm not a big fan of OpenACC, because it requires the user to make assumptions on the underlying hardware and quite a bit of thinking to get high-performance code. It is quite time consuming to integrate the pragmas so that the code performs. OpenACC capable compilers used to be very unstable, we regularly got compiler breakages and regressions. It got a bit better since we send the code to the vendors, but we still see regressions from time to time. All (usable) OpenACC implementations are proprietary, so we have a vendor dependency. The PGI OpenACC binaries used to be pretty slow, but it is now almost up to par with the Cray binaries.
The successor model to COSMO, ICON, developed at DWD and MPI is also written in Fortran using OpenACC pragmas (and I think OpenMP pragmas because they plan to use it on XeonPhi). The code is interesting in the sense that they are using a icosahedral grid, which stands in contrast to the square grid COSMO uses: Data accesses are not straightforward. Keep in mind that Fortran does not have easy abstractions for data accesses outside of square grids, so you have to use a couple of macros/functions/etc.. to get to the fields you need. Disclaimer: I have not seen the code myself, but the DWD certainly knows how to write high-performance code, so I'm certain most of the code is optimized.
- gbrown_ 10y agoDo you work at MeteoSwiss perchance?
- __mp 10y agoYes
- gbrown_ 10y agoHa though so, it's a small world ;) The MeteoSwiss utilizes K80 GPU's in the system hosted at CSCS though right? Or were you saying you don't expect a shift to accelerator-based weather models as a general trend?
- __mp 10y agoYes, exactly: Piz Kesch and Escha. The code is currently also run on Piz Daint on Pascal. Piz Kesch/Escha is pretty special because the system has 16 GPUs per node which makes the MPI configuration pretty hairy compared to one GPU systems like daint. I meant it as a general trend. There's interest from the community to run the GPU model on their system but unfortunately the official COSMO code is not fully GPU ready yet.
- gbrown_ 10y agoCool, thanks for replying. Didn't know you guys were actually running anything on Daint (I'd heard tests were planned). Your earlier comment was the first time I'd heard of the ICON model too. Quite an informative day on HN.
- __mp 10y agoYou're welcome. This is what they presented last year: ftp://ftp-anon.dwd.de/pub/DWD/Forschung_und_Entwicklung/CUS2016_presentations_PDF/06_Dynamics_and_Numerics/01_Invited_Z%E4ngl/ICON_CUS_2016.pdf very nice stuff! (there's a seminar at DWD next week where they probably will give an update). The ICON model is also interesting because they are able to interleave a local area model over Europe with their global model. Well, actually it's actually mostly Jenkins that runs the code for us on Daint. ETH, which I'm also affiliated with, is starting to use the model in GPU mode for climate runs on Daint. What are you working on at Cray?
- m_mueller 10y agoHi! I'm actually working on an OSS transpiler project that is made to convert CPU-only array based Fortran applications such that they can run on both CPU and GPU. It's been in the works since 2012 and I'm just now gearing up for a 1500x1300 2km run on a Japanese supercomputer, results will get submitted soon. https://github.com/muellermichel/Hybrid-Fortran https://github.com/muellermichel/Hybrid-Fortran
- __mp 10y agoOh that's pretty neat. A colleague of mine is working on a similar project applying directives to allow loop reorderings for OpenACC and OpenMP code (CLAW project). He's using the Omni-Compiler tough.