4 ms·
Only armchair spectators are still talking about OpenCL. It's as dead as disco.
by moonbug22 9y ago
Only armchair spectators are still talking about OpenCL. It's as dead as disco.
- diab0lic 9y agoThat is more or less what I was getting at, while trying to remain polite on the subject. The truth is that none of the deep learning backends have real support for it and OSX supports an outdated version (IIRC). It has seen very little adoption outside of a handful of niche use cases.
- abiox 9y agoare there insurmountable technical constraints there, or could it be resuscitated?
- diab0lic 9y agoIt could be resuscitated but it is unlikely without the backing of a major player.
- Athas 9y agoOpenCL is a very sound design on a technical level. Unfortunately, it's mostly designed to be a good target for code generation/libraries, and is awkward to use directly. The problems are that everything is very explicit and spelled out, and the kernel language itself is based on C, not C++, and segmented from the rest of the code base. In practice, you use OpenCL API calls to submit strings containing code to the OpenCL backend, which then compiles it and hands you back an opaque compiled program that you can execute. This makes it hard to write modular and maintainable code with OpenCL directly. The issues go away if you use a good OpenCL frontend. PyOpenCL for Python goes a long way towards this, and is not really any more awkward than the corresponding PyCUDA, and higher-level languages that generate OpenCL code, like Lift[0] or Futhark[1] (tooting my own horn here), remove the awkwardness completely. [0]: http://www.lift-project.org/ http://www.lift-project.org/ [1]: https://futhark-lang.org https://futhark-lang.org
- pjmlp 9y agoI think that has been the major failure from Khronos, being stuck with C mentality for their APIs. OpenGL ES only took off thanks to gaming on the iPhone, and now is deprecated on Apple platforms. Vulkan still lives in a C world, and the semi-official C++ bindings only exist thanks NVidia. OpenCL waited too long to support C++, Fortran and providing an infrastructure for compiler writers to add GPU support to their own languages. And two years later the majority of drivers are not there yet.
- roel_v 9y agoKhronos' problem is that they want to do everything, or not do it at all. Sycl would be 'real' c++ support, way beyond the 'oo wrapper around c api' that already exists in opencl headers. Nvidia is more 'good enough? Ship it' which has shown itself many times in the past to be a dominant strategy. (As much as I hate to admit that).
- mellinoe 9y agoIt seems like a C API is a much better choice if you want your library to be callable from many different languages. What is a reasonable alternative without giving that up?
- pjmlp 9y agoA bytecode format, like what CUDA, Metal Compute and DirectX Compute use. Which Khronos finally adopted as SPIR, but the drivers aren't there yet. Regarding the other Khronos APIs, a C API is like being stuck in a PDP-11 world. Many mix the idea of C API with OS ABI, it only happens to be the same if the OS APIs are exposed as plain C. There many cases where this isn't like it, e.g. mainframes, mobile OSes, and most userspace on OS X (Obj-C runtime) and Windows (.NET and COM).
- mellinoe 9y agoD3D, Vulkan, OpenGL (optionally) etc. do use a bytecode format, but that only covers shader modules, which are a small part of the overall API. I'm not that familiar with CUDA, but it looks like a large portion of it is delivered as a C API, judging by the bindings I see online.
- roel_v 9y agoIt's purely political. If nvidia would have a team of 3-5 ppl working op their opencl drivers/tooling, opencl would be as good as cuda. Nvidia would be stupid to do that, of course. Hence the Khronos play to fold opencl into vulkan. We'll see how that plays out.
- joefourier 9y agoI still use OpenCL on the daily, it allows me to program for both NVidia and AMD GPUS simultaneously with the minimum amount of pain. And you can still use it for mobile GPUs and specialized chips like the Myriad. Do you have another open, cross-platform, widely compatible GPU programming framework to recommend?
- nicwilson 9y agoThere is DCompute[1] which is able to generate SPIR-V and PTX simultaneously from the same D source. See my comment to the parent. http://github.com/libmir/dcompute http://github.com/libmir/dcompute
- viewtransform 9y agoHIP allows developers to convert CUDA code to portable C++. The same source code can be compiled to run on NVIDIA or AMD GPUs https://github.com/ROCm-Developer-Tools/HIP https://github.com/ROCm-Developer-Tools/HIP
- pjmlp 9y agoWhat mobile GPUs? On iOS it is kind of deprecated and the way forward is Metal Compute. On Android it never happened, Google created their own Renderscript dialect instead.
- gcp 9y agoSame here. What else can you do to ship GPU acceleration to AMD and NVIDIA people? The alternatives recommended here aren't even serious IMHO. I'd rather switch to CUDA and wait till Intel/AMD sort out a REAL compatibility layer than deal with those. Unless I'm mistaken, HIP still requires a separate compile for either platform and what runtime do they expect end users to have exactly?! At least CUDA and OpenCL are integrated in the vendor drivers. Vulkan compute with SPIR-V seems to be the only real solution, but even that is still very early. Sill waiting for proper OpenCL 2.0 support in NVIDIA drivers :P
- freeone3000 9y agoYou simply don't ship. Enterprise deep learning doesn't ship their training code - large models are trained on purpose-designed, dedicated hardware. Hardware compatability doesn't matter, software does. Even that's flexible if it's significantly faster. (The models can be executed on low-powered, commodity CPUs. No need for any GPU there.)
- nicwilson 9y agoThere is no reason that you can't create both OpenCL (i.e. SPIR-V) and PTX from the same source. Thats what I do for [DCompute](https://github.com/libmir/dcompute https://github.com/libmir/dcompute). Agnosticism at the library level/CPU side is next on my list.
- roel_v 9y agoThe problem isn't getting it to run, the problem is getting similar performance. You need self-tweaking algos to get some level of performance portability using just opencl, let alone across two apis.
- nicwilson 9y agoIndeed, there is a very large amount of parameterisation you can do in D with ease. You can, if you have prior knowledge of the hardware that it will run on, create different versions of the kernels to exploit that (e.g. with template kernels, which we support directly unlike OpenCL C++). Do you mean adaptive algorithms or dynamic recompilation? And yes I do expect that cross API will be difficult, for both running and getting good perf. But it is not just the room at the high performance end of the spectrum that is important, but also the lower end that is stifled by the barrier to entry that would benefit from the extra compute.
- roel_v 9y agoI meant dynamic (re)compilation yes - like adapting the stride in your kernel when you vary work group size for OpenCL kernels. I mean, I have something like that (a primitive self-tuning pre-run calibration step) and I only vary work group-, local- and vector size; and that's already a massive pain in the ass. It's ok on toy kernels but once you move beyond that, it just seeps into all your kernel code, making it almost into a meta-language. And then I'm not even talking about differences like using image types vs arrays for data that is 2d by nature. I don't see how I can generalize that; I just write various versions of my kernels. Which doesn't scale very much, to put it mildly. My point is - was drawn to OpenCL for its 'portability' claim, and yes kernels will 'run' on various hardware, but with massive differences in speed. what good does that portability do me? My workloads (scientific computing, branch heavy) are different from the typical ML half- or single precision MulDiv()/linear algebra applications, so all the hand-tuned CUDA libraries aren't even my concern. The elephant in the room is that performance doesn't depend on this API or that; it's in how you tune your algorithm to the actual hardware you're running on. Which is the direct opposite of portability. To come back to the post I was replying to - yes it's 'trivial' (I mean, a lot of work, but technically not hard, not to belittle your work) to compile almost any statically typed code into either SPIR-V or PTX or any other future format for that matter. But that doesn't mean that it will work 'well' (not even 'optimal') on other hardware. In the real world, you're almost always better off just spending a few hundred to get the same GPU as whoever wrote the kernels tested them on (or if you're running your kernels on existing clusters, to focus on optimizing them for what you know you'll be running them on). Oh and all of that is just considering GPU's. I mean, when you read an OpenCL book they make it seem (in the first few chapters) that you don't even have to think about whether you'll be running on a CPU or a GPU. And then you accidentally use USE_HOST_PTR instead of COPY_HOST_PTR (or the other way around) in the wrong place, and all of a sudden your code is 10x slower than the sequential version of your algo even. What I'm saying is - I no longer believe in easy to use abstractions for these purposes. If you want speed, you code to the metal, and/or you tweak your abstractions to your specific use case. Yes this is a lot of work. And if you don't need speed, you just throw in a few std::thread's here and there and call it a day. But that's just my conclusion for my use cases.
- chadnickbok 9y agoDisco will never die.
- gcp 9y agoI'll continue using it as long as the users are still split between NVIDIA and AMD cards. If AMD GPUs die out, CUDA it is. Or maybe give up and use a wrapper library for whatever 10 alternatives-only-supported-by-one-marginal-vendor there are. (like Apple)