3 ms·
Seems like it currently supports only CUDA and ROCm. Are there plans to support any other GPU vendors, such as Intel? The example also seems a little bit high
by ColonelPhantom 3y ago
Seems like it currently supports only CUDA and ROCm. Are there plans to support any other GPU vendors, such as Intel?
The example also seems a little bit high on 'magic'. Chapel is told to execute on GPU, yet it somehow decides "this loop is order-dependent so it should run on the CPU anyway"? It's not a bad approach necessarily, but I don't think you can get any such implicit serialization when directly programming CUDA C++ or similar. (I think Futhark, another language for high-level parallel computing, also doesn't suffer from this kind of thing, thanks to being purely functional.)
- e-kayrakli 3y agoRe Intel support: That's definitely in our plans. However, there are also many other areas that we are actively working on to add more features, fix bugs and improve performance. When prioritizing we typically make decisions based on what our current and potential users might need in the language. Frankly, we are not seeing a big push for Intel GPU support so far. So, currently it is not near the top of our priorities. If you (or other readers) have any input on that matter where lack of Intel support might be a blocker for testing Chapel and/or its GPU support out, definitely let us know. Re implicit serialization: To clarify; the serialization based on order-dependence is not implicit. The users should use `for` loop if their loop is order-dependent and `foreach` (and `forall`) if their loop is order-independent. In other words, the Chapel compiler doesn't make decisions about order-dependence. In particular for GPU execution: a `for` loop will never turn into a GPU kernel. There are however some cases where a `foreach` does not turn into a kernel. You may be referring to those cases, but that's not related to order-dependence. Some Chapel features cannot execute on GPU. If your `foreach` loop's body uses any of those features then it will not be launched as a kernel even though `foreach` signals order-independence. Now, a subset of such features that makes an order-independent loop GPU-ineligible are there because we haven't gotten a chance to properly address them, yet. Another subset of such features will remain thwarters for a longer time and maybe forever. For example, your `foreach` loop could be calling an extern host function.
- adastra22 3y agoWhat about Apple Silicon / Metal support?
- e-kayrakli 3y agoThanks for bringing this up. I posted an answer to the same question here: https://news.ycombinator.com/item?id=39009566 https://news.ycombinator.com/item?id=39009566 But seeing that we already have two questions about it makes me think whether this should be something we should think more about when we are prioritizing work.
- adastra22 3y agoThis so going to sound extreme, but it is true: I personally and professionally won’t use it for anything until it has Metal support. The simple fact is that me and my team do our development largely on MacBook Pros and Mac Studios. We have some GPU rigs for running production code on Lunux or Windows with NVIDIA GPUs, but all of our developer tooling is on macOS. Anything that can’t run natively in recent Apple hardware gets second-class support.
- bradcray 3y agoThat doesn't seem extreme to me, as I generally feel similarly. If you (or other readers) are genuinely interested in using Chapel with Metal, please open an issue on our GitHub repository capturing your request, as that would be valuable to us. Just to make sure it didn’t get lost, note that it is possible to develop GPU code in Chapel on a MacBook using the cpu-as-device mode Engin mentions above, and then deploy it on NVIDIA GPUs on production systems by recompiling. This is how I develop/debug GPU computations in Chapel.
- misnome 3y ago> If your `foreach` loop's body uses any of those features then it will not be launched as a kernel even though `foreach` signals order-independence. Is this signalled/warned about so that you don’t accidentally use one of these features and kill your performance? Or a way to indicate that you specifically intend it to be run on GPU?
- e-kayrakli 3y ago[dupe]
- bradcray 3y ago@ColonelPhantom: Thanks very much for your questions. The following are answers I'm relaying from Engin Kayraklioglu, who heads up the Chapel GPU effort: Re Intel support: That's definitely in our plans. However, there are also many other areas where we are actively working on to add more features, fix bugs, and improve performance. When prioritizing, we typically make decisions based on what our current and potential users might need in the language. Frankly, we are not seeing a big push for Intel GPU support so far. So, currently it is not near the top of our priorities. If you (or other readers) have any input on that matter where lack of Intel support might be a blocker for testing Chapel and/or its GPU support out, definitely let us know. Re implicit serialization: To clarify; the serialization based on order-dependence is not implicit. The users should use a `for` loop if their loop is order-dependent and `foreach` (and `forall`) if their loop is order-independent. In other words, the Chapel compiler doesn't make decisions about order-dependence. In particular, for GPU execution a `for` loop will never turn into a GPU kernel. There are, however, some cases where a `foreach` does not turn into a kernel. You may be referring to those cases, but that's not related to order-dependence. Some Chapel features cannot execute on a GPU. If your `foreach` loop's body uses any of those features then it will not be launched as a kernel even though `foreach` signals order-independence. Now, a subset of such features that makes an order-independent loop GPU-ineligible are there because we haven't gotten a chance to properly address them, yet. Another subset of such features will remain thwarters for a longer time and maybe forever. For example, your `foreach` loop could be calling an external host function.
- bradcray 3y agoSorry for what now appears to be a double-post. Engin had just registered for HN, hadn't seen his reply going through, so asked me to relay it. Re-reading this Q+A this morning, I also wanted to clarify one thing, which is that when a 'foreach' or 'forall' does end up being executed on the CPU, that doesn't mean it has been serialized. 'foreach' loops on the CPU are candidates for vectorization while 'forall' loops typically result in multicore task-parallelism with each task also being a candidate for vectorization.