3 ms·
TL;DR: the parent doesn't know what they are talking about and for some reason is burnt over OpenCL which is a horrible ""standard"" (Khronos extension over C a
by volta83 5y ago
TL;DR: the parent doesn't know what they are talking about and for some reason is burnt over OpenCL which is a horrible ""standard"" (Khronos extension over C and C++, not actual C++ support, in the C++ ISO standard proper).
Maybe the reason people like HIP and OneAPI is because they are not aware that you can program GPUs using ISO C++ and ISO Fortran because... Intel and AMD don't support those?
> Absolutely false. [...] If you mean something like `std::par` [...] not really GPU-specific
std::par is the ISO C++ standard feature for parallel programming, this includes multi-core CPUs, SIMD, GPUs, FPGAs, etc.
NVIDIA compilers run std::par on both CPUs and GPUs.
I've been using it for over a year, and so have a lot of people I work with.
You claim that this is false, but this is demonstrably true, and there are many peer-reviewed papers of people using this with NVIDIA compilers already.
The claim that this model of GPU programming is not portable, is false as well, since Intel and AMD claimed that this is portable to their GPUs when they added it to the C++ standard in 2017.
The claim that this does not deliver performance portability is also false, since we have verified performance portability across 3 NVIDIA GPU architectures, and using Kokkos, which exposes a similar programming model, also to Intel and AMD architectures.
- jeeceebees 5y agoHow does the performance between GPU programs written with std::par compare to those written in CUDA? Do you happen to know of any online resources that show a comparison of the kernel code and performance of the two frameworks on common tasks?
- volta83 5y agoThis paper ported a CFD application, which had a tuned CUDA implementation, to std::par: https://arxiv.org/pdf/2010.11751.pdf https://arxiv.org/pdf/2010.11751.pdf . In Table 3, first and last columns shows the performance of CUDA and std::par in % of theoretical peak. The rows show results for different GPU architectures. On V100, CUDA achieves 62% theoretical peak and std::par 58%. The amount of developer effort required to achieve over 50% theoretical peak with std::par makes it a no brainer IMO. If there is one kernel where you need more performance, you can always implement that kernel in CUDA, but for 99% of the kernels in your program your time might be better spent elsewhere.