4 ms·
2-3 most fundamental parallel algorithms Would you be able to list them? I'm asking because I do develop programming languages for parallel hardware. It w
by 9q9 6y ago
2-3 most fundamental
parallel algorithms
Would you be able to list them? I'm asking because I do develop programming languages for parallel hardware. It would be very useful for me to look at your examples and apply them to my PL design ideas.
An alternative way I might be rendering my question could be: what primitives do you recommend a language should provide that's currently missing in the languages you evaluate?
- fluffything 6y ago> Would you be able to list them? scan (e.g. prefix sum), merge, partition, reduce (e.g. minimum), tree contraction, sort > what primitives do you recommend a language should provide that's currently missing in the languages you evaluate? Hardware primitives and powerful capabilities for defining safe abstractions over those. Chances are that, e.g., I need to implement my own "primitive-like" operation (e.g. my own sort or scan). If your language provides a "partition" primitive, but not the tools to implement my own, I can't use your language. When people create languages for GPUs, for some reason they add the C++ STL or similar as primitives, instead of providing users the tools to write such libraries themselves. The consequence is that those languages end up not being practical to use, and nobody uses them.
- 9q9 6y agoThanks. How to abstract "hardware primitives" in a way that can be instantiated to many GPU architectures without performance penalty and be useful for higher-level programming is not so clear. How would you, to take an example from the CPU world, fruitfully abstract the CPU's memory model? As far as I'm aware that's not a solved problem in April 2020, and write papers on this subject are still appearing in top conferences .
- deleted 6y ago[deleted]
- fluffything 6y agoThere exists a language that can optimally target AMD's ROCm, NVIDIA's PTX, and multicore x86-64 CPUs: CUDA C++. This language has 3 production quality compilers for these targets (nvcc, hip, and pgi). Unfortunately, this language is from 2007, and the state-of-the art has not really been improved since then (except for the evolution of CUDA proper). It would be cool for people to work on improving the state of the art, providing languages that are simpler, perform better, or are more high-level than CUDA, without sacrificing anything (i.e. true improvements). There have been some notable experiments in this regard, e.g., Sequoia for roadrunner scratchpads had very interesting abstraction capabilities over the cache hierarchy, that could have led to a net total improvement over CUDA __shared__ memory. Most newer languages have the right goal of trying to simplify CUDA, but they end up picking trade-offs that only allow them to do so by sacrificing a lot of performance. That's not an interesting proposition for most CUDA developers - the reason they pick up CUDA is performance, and sacrificing ~5% might be acceptable if the productivity gains are there, but a 20-30% perf loss isn't acceptable - too much money involved. One can simplify CUDA while retaining performance by restricting a language to GPUs and a particular domain, e.g., computer graphics / gfx shaders or even stencil codes or map-reduce, etc. However, a lot of the widely advertised languages try to (1) target both GPUs and CPUs, compromising on some minimum common denominator of features, (2) support general-purpose compute kernels, and (3) significantly simplify CUDA, often doing so by just removing the explicit memory transfers, which compromises performance. These languages are quite neat, but they aren't really practical, because they aren't really true improvements over CUDA. Most companies I work with that use CUDA today are using the state of the art (C++17 or 20 extensions), experimental libraries, drivers, etc. So it isn't hard to sell them into a new technology, _if it is better than what they are already using_.
- 9q9 6y agoThanks. [1] contains an interesting discussion of Sequoia, and why it was considered a failed experiment, leading to Legion ][2], its successor. [1] https://news.ycombinator.com/item?id=18009581 https://news.ycombinator.com/item?id=18009581 [2] https://legion.stanford.edu/ https://legion.stanford.edu/
- 6y ago