3 ms·
The problem isn't that a particular library is AMD or NVIDIA only in terms of computation, but NVIDIA or AMD only in terms of performance. GPU implementations a
by 14113 9y ago
The problem isn't that a particular library is AMD or NVIDIA only in terms of computation, but NVIDIA or AMD only in terms of performance. GPU implementations are usually optimised for a specific architecture, which is then the "supported" architecture.
On occasion, such implementations do use vendor specific tools (such as CUDA), but there are a plethora of tools such as OpenCL, SyCL etc that provide portability - but not always performance portability, meaning that they will still be tuned to a specific architecture.
For performance portability, the LIFT project (http://www.lift-project.org/ http://www.lift-project.org/) provides a partial solution. Our approach relies on a high level model of computation (think or something like a functional, or pattern based programming language) coupled with a rewrite-based compiler that explores the space of OpenCL programs with which to implement a computation.
That lets us "optimise" a given implementation to a specific architecture, entirely automatically, in a way that many other low level approaches simply aren't able to, as they contain too many implementation (rather than computation) details.
- marmaduke 9y agoLooks interesting conceptually but it would be good to refute the "sufficiently clever compiler" in anticipation. My beef with this sort of thing or Furthark or OpenACC is that they don't seem to understand complex data layouts required for real problems.
- 14113 9y agoDefine "complex data layouts"? One reason that tools like Futhark often don't is that complex data structures don't provide good performance on GPUs. If, however, you mean complicated compositions of arrays, then that is something we support, as well as efficient ways for describing (e.g.) coalesced accesses or stencil operations.
- marmaduke 9y agoThanks for the reply. I should not have included Furthark with OpenACC. The layouts I have in mind are ring buffer of arrays and ND arrays. I think Furthark would work well actually but I simply had not the time to get accustomed to it.