4 ms·
The author says he's happy with the performance improvements he made, but--even at the size where the Rust version performs best--he gets much less than a 2x sp
by mindB 6y ago
The author says he's happy with the performance improvements he made, but--even at the size where the Rust version performs best--he gets much less than a 2x speedup. After a couple months of effort, that's not a whole lot to show. This reads like a case study in why it's probably a bad idea to rewrite even when you think you have compelling reason for it.
- pjmlp 6y agoAlso the author has not used state of the art compilers like XL, Intel and PGI.
- milancurcic 6y agoIndeed, and this depends between applications. For the WRF (Weather Research and Forecasting) model, I get 3-4x speed up with Intel Fortran compiler over gfortran. I don't see where the Author mentions the Fortran compiler used, though.
- wycy 6y ago> I get 3-4x speed up with Intel Fortran compiler over gfortran Wow, even at similar optimization levels? -O3 for each?
- milancurcic 6y agoNot quite, but at optimization levels set by default WRF configuration for each compiler (definitely not a fair comparison): gfortran: -O2 -ftree-vectorize -funroll-loops ifort: -O3 I don't have the timing results anymore. This was in 2018 on Xeon Platinum 8168. I recently tried replacing -O2 with -Ofast -ffast-math to gfortran settings, which gives about 18% speed up. So still far from Intel. I recently proposed it here [1]. [1] https://github.com/wrf-model/WRF/issues/1254 https://github.com/wrf-model/WRF/issues/1254
- RockIslandLine 6y agoHave you posted about this on the gfortran mailing lists? They are generally interested in examples that show where the compiler could be improved.
- mkbosmans 6y agoI doubt they are short on inspiration for further improvements. A big part of the speed advantage of ifort is the aggressive loop unrolling, pipelining, splitting and multiversioning. Coming up with these transformations is not the difficult part. That would be actually implementing these transformations correctly and keep the whole compiler optimization framework maintainable. And of course there is certainly a cost for the improved runtime. Primarily in compile time and code size. For the Intel compiler it generally makes sense to sacrifice these in favor of runtime performance, because it is used on these kind of scientific codes a lot, but for gfortran the balance might different.