4 ms·
I strongly dislike CUDA. Once you have allowed that proprietary cr*p into your C++ codebase, it is very hard to get rid, and you end up with code that is either
by jacobgorm 17d ago
I strongly dislike CUDA. Once you have allowed that proprietary cr*p into your C++ codebase, it is very hard to get rid, and you end up with code that is either tied to a single vendor or an #ifdef hell, probably both.
The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually, like in Metal, OpenCL, and D3D12, etc.
These days we even have DSLs like Triton that make kernel writing much more ergonomic than anything you would hope to achieve in Rust.
- bigyabai 17d agoIs this satire? D3D12 and Metal aren't any less proprietary than CUDA.
- deleted 17d ago[deleted]
- jacobgorm 17d agoYou can call their APIs without needing to compile your code with a proprietary compiler or adopt a bastardized version of C++.
- bigyabai 17d agoSounds like a C problem, not a CUDA problem.
- pjmlp 16d agoActually no. You need Objective-C, Swift, and Metal is a C++14 dialect with extensions. You may refer to the C++ bindings, which still not obviate the need for the C++14 dialect in the shaders, and it only works, because there is a shim to call the Objective-C runtime from C++. Likewise there are DirectX COM interfaces that are really only usable from Visual C++ COM extensions, and the HLSL semantics depend very much on which compiler is being used, hence why there is finally a language reboot going on.
- pavon 17d ago> The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually Isn't that how CUDA code is normally written?
- jacobgorm 17d agoNo. CUDA allows you to write all the code in a single file, and uses a preprocessor to split it back out and pass it through separate compilers, one for host and one for device.
- compiler-guy 17d agoThis true, but you can write the two separately if you want. The disadvantages of writing them together are listed in the various parent posts. But some code authors really like the convenience of having the two in the same file.
- melodyogonna 17d agoYou could also use Mojo, one language for all targets.
- carefree-bob 17d agoI began to lose interest after the acquisition. Have you been following along, are they still going to open source it?
- YuechenLi 17d agoI thought they already did and released the compiler source code under Apache 2.0.
- ecl3ctic 17d agoThe Mojo compiler has been open source for over a month now. And the Mojo standard library has been open source for over a year. It’s all open source. Go check it out!
- carefree-bob 17d agoNice, thank you. There is an old python project I've been thinking about converting to Mojo.
- adgjlsfhk1 17d agoOr julia if you want a much more mature ecosystem.
- patagurbon 17d agoI highly recommend Julia for (scientific) GPU programming but it would be nice if there was a larger community and/or funding behind the GPU side of things. It has very few core devs for what it is.
- eggy 17d agoJulia has had a great CUDA story for a few years now, and this about 9 days old. Rust rejects buffer aliasing at compile time using Rust's borrow checker, but shared memory in cuda-oxide currently requires unsafe, but then there's HuggingFace's Grout and mistral.rs, so yeah, Rust is picking up ground here on Julia. How is OpenCL's performance these days?
- fg137 17d ago> Once you have allowed that proprietary cr*p into your C++ codebase People have been doing that all the time for every kind of codebase. It's just part of the business. I don't see how it's worth having any emotions or opinions about it. Seems like you are wasting your energy. Are win32 APIs proprietary? So you decide to use them, use a wrapper/UI framework, or don't develop for Windows. Easy choice. Developing for embedded devices? So you read the manufacturers manual and implement based on the spec, use some sort of HAL if they are available, or you don't have a job. Even simpler.
- jacobgorm 17d agoCUDA is not an API, CUDA is a language, so you cannot make that comparison.
- esseph 17d ago> The CUDA runtime is a special case of one of the libraries provided by the CUDA Toolkit. The CUDA runtime provides both an API and some language extensions to handle common tasks such as allocating memory, copying data between GPUs and other GPUs or CPUs, and launching kernels. The API components of the CUDA runtime are referred to as the CUDA runtime API. From: https://docs.nvidia.com/cuda/cuda-programming-guide/01-introduction/cuda-platform.html https://docs.nvidia.com/cuda/cuda-programming-guide/01-intro...
- pjmlp 17d agoCUDA is neither an API, nor a language, it is an ecosystem.
- fc417fc802 17d agoThat's a nice way of saying that it's a dependency clusterfuck. I've never understood why we can't just expose the GPU ISA directly the way the CPU does. It's all getting compiled down at the end of the day so someone has to write a compiler for it either way. We'd be substantially better off IMO if it was all built directly into LLVM and then let middleware sort out the details.
- cpill 17d agoyeah, just write a stub/wrapper around it and abstract. it's the classic coupling problem. nothing to do with CUDA
- tombert 17d ago> I strongly dislike CUDA. Once you have allowed that proprietary cr*p Genuine question...why not just type "crap"? It's not even that much of a curse, but I've never really understood the point of self-censorship. If you don't want to curse then you could just use a non-curse word.
- lovelearning 17d agoIt may be to bypass censorship, rather than self-censorship. Some platforms block or shadowban comments with curse words. Not sure about this platform.
- arcanemachiner 17d agoHN definitely doesn't give a crap about that word.
- smnplk 17d agocan confirm, looks like crap is not on a list
- flamedoge 17d agopretty crappy list
- tombert 17d agoI have written many words far worse than "crap" on this site. I haven't gotten in trouble over it yet. I do find it a little amusing, because commenters stopped criticizing my cursing the moment I started getting a good chunk of karma here. I remember in 2016 someone criticized me for using the term "shitposting"...I don't think I've gotten that kind of criticism since 2016 though.
- protocolture 17d ago[dead]
- nicwilson 17d agoLaunching kernels manually is an error prone PITA which I believe is the principle reason for CUDA's popularity. Having the compiler give an error when you mess up is a huge benefit. But having the compiler allow you to express "I want to launch this kernel over a grid with these dimensions, with these arguments" as a single expression is where the vast majority of the value comes from. The having it all in a single file is mostly an artefact of the fact that it is C++, because C++ is single file at a time compilation. In D (which is multiple files in a single compiler invocation) with DCompute (which targets CUDA and OpenCL with upcoming support for Vulkan and Metal), you are required to write the kernels in a separate module, but you get all the benefits of the compiler complaining when you mess up _and_ the expressivity of "launch me this kernel".
- oblio 17d ago> Having the compiler give an error when you mess up is a huge benefit. Shouldn't this be alleviated by the current code generation machines?
- nicwilson 17d agoWell yeah, but then you are using code generation, not writing code directly.
- oblio 17d agoI meant LLMs :-)
- high_na_euv 17d agoYou are trying to say that llm can replace compiler?
- oblio 17d agoIn general, no? But they should help with this part: > Launching kernels manually is an error prone PITA which I believe is the principle reason for CUDA's popularity.
- winwang 17d agoHaving also played with Metal and WebGPU (at least years ago), I would say that CUDA is, amazingly, the best GPGPU API we have. Do I wish we had an open source parallel programming language as good or better than it? Yes. But asymmetrically hating on CUDA like this is how we continue to lag behind it in UX. > The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually Not to mention that this is a completely sane way to use CUDA as well.
- pjmlp 17d agoPeople that attack proprietary APIs always miss the point why most devs outside FOSS circles prefer them. Turns out when one isn't ideologically against something they aren't willing to put up with a lesser experience just for the cause.
- darkwater 17d agoI know it's not the same thing because proprietary vs open software it's way less important but, generally if you are not ideologically against something you can easily follow the stream and do lot of nefarious actions, especially if the action has enough degrees of separations from the actual nefast outcome.
- 15155 17d agoI don't mind CUDA, I do mind that all of the SDKs don't dynamically load the various CUDA shared libraries at runtime.. intertwining itself into your application linking process makes for extreme binary portability inconvenience.
- uncle_kostya 16d agoThere a flavor of CUDA runtime libraries that binds at runtime, so you can have a single binary that runs with CUDA and without it. I did this at work. Obviously you need to check if CUDA is available before trying to execute kernels, or it will error out.
- 15155 16d agoSure, I've written them. None of the NVIDIA-provided SDKs are like this, including the new Rust one.
- anon291 17d ago? I find it hard to see the issue here. Just put it in a separate file and call it?
- throwaway334212 17d agoAnyone here looking at Modular's offerings?
- bsaul 17d agoi'm surprised modular's doesn't get much traction. The promise seems super interesting, and chris latner has the record to back up his claims. If someone has an explanation..
- harrison_clarke 17d agofrom what i can tell, you're going to be stuck with that no matter what you do i'm currently using vulkan, and HLSL via dxc. which should be portable but it's not. apple refuses to support vulkan, and relies on moltenvk and there's a bunch of OS/hardware/driver differences no matter what you do, that you'll probably have to feature test for, and compile a few different versions of your code no matter what you do i think if you're doing something that you don't have to distribute to customers, just picking one stack and getting locked in has some appeal. it leaves you vulnerable to lockin. but, especially in the age of ai, "claude, port this to vulkan" seems like a good enough defense against that
- unPeuResilient 17d ago[dead]
- wangxili1997 17d ago[flagged]
- andirk 17d agoWhat is a proprietary crop?
- mstkllah 17d agoIt's actually creep.
- mschuetz 17d ago> The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually, Yes, I also prefer doing it that way, but in Cuda with the driver API. Allows you to handle kernels like shaders, including editing and hot-reloading at runtime. The reason I'm sticking with CUDA is because it's by far the most convenient API to use, without nonsense like 50-liners to alloc memory or the need to manage descriptors, bindings, queue families, etc.
- david-gpu 17d ago> The reason I'm sticking with CUDA is because it's by far the most convenient API to use, without nonsense like 50-liners to alloc memory or the need to manage descriptors, bindings, queue families, etc. I was there when the OpenCL committee was deciding on that sort of stuff. As I recall, and it's been two decades and a lot of sleepless nights since then, there was real pushback at the time against OpenGL-style default bindings. So folks didn't want to establish an implicit command queue or any other default objects attached to other objects. Part of it is because OpenGL was perceived as clumsy and passé, some of it was because it is not friendly to multi-threaded applications. Those first meetings were a shitshow full of tension, implicit threats from Apple, and backroom deals. Kudos to Neil Trevett for chairing the group; I I bet it wasn't fun for him either.
- mschuetz 17d agoThat's unfortunate. Cuda has shown that, when done right, defaults and a convenience layer can make for a well received API without sacrificing performance.
- david-gpu 17d agoYes, I wanted defaults as well, particularly a default context and command queue. Design by committee is a real phenomenon. And people in a committee know that, but they are also helpless.
- ActorNightly 17d agoYep. Ive essentially followed that paradigm with Python and C. I start out writing Python code. If I need something to run fast, I build a standalone C application that either reads from a file or listens on a socket, and just invoke it from Python. No need to write the entire thing in Rust and deal with all its semantics when it will be at best like 2% faster.
- melihelibol 16d agoYou don't need to use the CUDA (SIMT) programming model if you don't like it. The project includes cutile, which lets you program the GPU using tensors. It feels a lot like programming the GPU using numpy and triton.
- jacobgorm 16d agoAs it happens, I just got my employer's permission to release as open source a Triton back-end for Metal and D3D12 GPUs here: https://github.com/dropbox/neso https://github.com/dropbox/neso . As an example of you how can use it to deploy real models there is this project doing ASR and TTS: https://github.com/dropbox/nspeech https://github.com/dropbox/nspeech . Finally, I am also going to be switching the inferencing part of Witchcraft from current Candle on MacOS and OpenVINO on Windows to just Candle with Neso; https://github.com/dropbox/witchcraft https://github.com/dropbox/witchcraft
- bbkane 16d agoThat sounds like it'll be easier to maintain. Will it also be faster?
- jacobgorm 16d agoIt is currently faster than the stock Candle / MPS shaders it replaces on MacOS/ARM64, and IIRC a bit slower than OpenVINO/CPU on my old Windows laptop, where I never got OpenVINO/GPU to compute correctly. Candle didn't have support for GPUs on MacOS/Intel, and OpenVINO ceased to be supported there. Compared to OpenVINO (I tried ONNX runtime too, but never got it produce correct outputs with my quantized models) it is very nice to be able to build the exact kernels I need, at the quantization settings and precision that works for the models I have and with the custom operators required (speech models do a lot of non-standard stuff), run from a single set of sources, and not have to ship a hefty third-party DLL, and having to deal with their memory leaks and other stability issues.
- orangelimetea 16d ago[dead]
- lemonlimesoda 16d ago[dead]
- kocsonya 14d agoKeep in mind for the future there is mojo out there and it's open-source. I love it!