3 ms·
To be fair the entry barrier got a lot lower over the past ~5 years. Now you can write very good CUDA kernels with just a few lines of python DSL code. Zero cpp
by KeplerBoy 2mo ago
To be fair the entry barrier got a lot lower over the past ~5 years. Now you can write very good CUDA kernels with just a few lines of python DSL code. Zero cpp boilerplate and zero explicit compiler calls.
Stuff like Triton, nvidia warp (the language), numba, cupy jax/pallas and so many others really paved the way. You can start out really high-level, run a profiler and then dive deep into the bottlenecks.
TL,DR: Keep going, it's a great time to have fun with GPUs.
- embedding-shape 2mo ago> To be fair the entry barrier got a lot lower over the past ~5 years. Now you can write very good CUDA kernels with just a few lines of python DSL code. Zero cpp boilerplate and zero explicit compiler calls. Well, yeah, but what I've being doing is learning proper CUDA, not "Python-compiled-to-CUDA" (otherwise it'd take like a just a week to understand enough :P ) and that's looking more or less the same today (although bunch of more complicated stuff piled on top of the fundamentals) as it used to, AFAIK. With that said, the environment is a lot simpler to setup today at least :)
- KeplerBoy 2mo agoI wouldn't call one proper CUDA and the other one some dumbed down version. Nvidia really seems to be pushing for these DSLs to be first class within the ecosystem. In some cases probably even more cutting edge than the nvcc frontend, since it's easier to do some experimenting on a new niche package than on the tool everyone relies on. I believe more and more production code is running kernels which didn't originate from the traditional cuda cpp route.
- embedding-shape 2mo ago> I wouldn't call one proper CUDA and the other one some dumbed down version. Nvidia really seems to be pushing for these DSLs to be first class within the ecosystem. I wouldn't say one is dumbed down either, just different, at least the entrypoints and how you end up using the different solutions. I'm currently experimenting with cuda-oxide for some new simulations, and managed to keep the entire simulation within just Rust essentially, while going the "traditional" (maybe better term than "proper"?) way I've ended up with a bunch of .cu files and then integrating them (via cudarc usually). Kernels themselves feel the same across both, but the integration clearly makes them different enough that I think it's worth distinguishing them, at least for clarity if nothing else. If someone else already knew Rust but not C++, wanted to get into CUDA programming, going the cuda-oxide route would probably be easier and more familiar, than cudarc, I'd guess. Personally I'm not sure what route I prefer yet, both (as always?) have tradeoffs.