14 ms·
Nvidia adds native Python support to CUDA
- jbs789 1y agoSource?
- macksd 1y agoIt seems to pretty much be about these packages: https://nvidia.github.io/cuda-python/cuda-core/latest/ https://nvidia.github.io/cuda-python/cuda-core/latest/ https://developer.nvidia.com/nvmath-python https://developer.nvidia.com/nvmath-python
- dachworker 1y agoMore informative than the article: https://xcancel.com/blelbach/status/1902113767066103949 https://xcancel.com/blelbach/status/1902113767066103949 cuTile seems to be the NVIDIA answer to OpenAI Triton.
- pjmlp 1y agoThe plethora of packages, including DSLs for compute and MLIR. https://developer.nvidia.com/how-to-cuda-python https://developer.nvidia.com/how-to-cuda-python https://cupy.dev/ https://cupy.dev/ And "Zero to Hero: Programming Nvidia Hopper Tensor Core with MLIR's NVGPU Dialect" from 2024 EuroLLVM. https://www.youtube.com/watch?v=V3Q9IjsgXvA https://www.youtube.com/watch?v=V3Q9IjsgXvA
- diggan 1y agoI'm no GPU programmer, but seems easy to use even for someone like me. I pulled together a quick demo of using the GPU vs the CPU, based on what I could find (https://gist.github.com/victorb/452a55dbcf59b3cbf84efd8c3097a855 https://gist.github.com/victorb/452a55dbcf59b3cbf84efd8c3097...) which gave these results (after downloading 2.6GB of dependencies of course): Creating 100 random matrices of size 5000x5000 on CPU... Adding matrices using CPU... CPU matrix addition completed in 0.6541 seconds CPU result matrix shape: (5000, 5000) Creating 100 random matrices of size 5000x5000 on GPU... Adding matrices using GPU... GPU matrix addition completed in 0.1480 seconds GPU result matrix shape: (5000, 5000) Definitely worth digging into more, as the API is really simple to use, at least for basic things like these. CUDA programming seems like a big chore without something higher level like this.
- deleted 1y ago[deleted]
- rahimnathwani 1y agoThank you. I scrolled up and down the article hoping they included a code sample.
- diggan 1y agoYeah, I figured I wasn't alone in doing just that :)
- rahimnathwani 1y agoEDIT: Just realized the code doesn't seem to be using the GPU for the addition.
- wiredfool 1y agoCurious what the timing would be if it included the memory transfer time, e.g. matricies = [np.random(...) for _ in range] time_start = time.time() cp_matricies = [cp.array(m) for m in matrices] add_(cp_matricies) sync time_end = time.time()
- hnuser123456 1y agoI think it does?: (the comment is in the original source) print("Adding matrices using GPU...") start_time = time.time() gpu_result = add_matrices(gpu_matrices) cp.cuda.get_current_stream().synchronize() # Not 100% sure what this does elapsed_time = time.time() - start_time I was going to ask, any CUDA professionals who want to give a crash course on what us python guys will need to know?
- apbytes 1y agoWhen you call a cuda method, it is launched asynchronously. That is the function queues it up for execution on gpu and returns. So if you need to wait for an op to finish, you need to `synchronize` as shown above. `get_current_stream` because the queue mentioned above is actually called stream in cuda. If you want to run many independent ops concurrently, you can use several streams. Benchmarking is one use case for synchronize. Another would be if you let's say run two independent ops in different streams and need to combine their results. Btw, if you work with pytorch, when ops are run on gpu, they are launched in background. If you want to bench torch models on gpu, they also provide a sync api.
- aixpert 1y agothank God, Pytorch gained so much momentum before this came out, Now we have a true platform independent semi standard For parallel computations. We are not stuck with NVIDIA specifics. It's great that parts of pie torch which concern the NVIDIA backend can now be implemented in Python directly, The important part that it doesn't really matter or shouldn't matter for end users / Developers that being said, maybe this new platform will extend the whole concept of on GPU computation via Python to even more domains like maybe games. Imagine running rust the Game performantly mainly on the GPU via Python
- disgruntledphd2 1y agoThis just makes it much, much easier for people to build numeric stuff on GPU, which is great. I'm totally with you that it's better that this took so long, so we have things like PyTorch abstracting most of this away, but I'm looking forward to (in my non-existent free time :/ ) playing with this.
- wafngar 1y agoWhy not use torch.compile()?
- the__alchemist 1y agoRust support next? RN I am manually [de]serializing my data structures as byte arrays to/from the kernels. It would be nice to have truly shared data structures like CUDA gives you in C++!
- taminka 1y agoeven putting aside how rust ownership semantics map poorly onto gpu programming, ml researchers will never learn rust, this will never ever happen...
- the__alchemist 1y agoGPGPU programming != ML.
- pjmlp 1y agoWhile I agree in principle, CUDA is more than only AI, as people keep forgetting.
- malcolmgreaves 1y agoML reachers don’t write code, they ask ChatGPT to make a horribly inefficient, non-portable notebook that has to be rewritten from scratch :)
- staunton 1y ago
- gymbeaux 1y agoThis is huge. Anyone who was considering AMD + ROCm as an alternative to NVIDIA in the AI space isn’t anymore. I’m one of those people who can’t (won’t) learn C++ to the extent required to effectively write code for GPU execution…. But to have a direct pipeline to the GPU via Python. Wow. The efficiency implications are huge, not just for Python libraries like PyTorch, but also anything we write that runs on an NVIDIA GPU. I love seeing anything that improves efficiency because we are constantly hearing about how many nuclear power plants OpenAI and Google are going to need to power all their GPUs.
- ErrorNoBrain 1y agoThey are, if they cant find an nvidia card
- pjmlp 1y agoNVidia cards are everywhere, the biggest difference to AMD is that even my lousy laptop GeForce cards can be used for CUDA. No need for a RTX for learning and getting into CUDA programming.
- gymbeaux 1y agoTrue, although I believe Maxwell is the oldest supported architecture for the current CUDA 12.x. Maxwell (eg GTX 980) came out around 2013, if memory serves. 10+ years of support is not bad at all considering ROCm supports only like 3 consumer AMD GPUs. So your lousy laptop GTX 750Ti ehhh probably can’t practically be used for CUDA. But your lousy 1050Ti Max-Q? Sure.
- ferguess_k 1y agoJust curious why can't AMD do the same thing?
- bigyabai 1y agoIt can be argued that they already did. AMD and Apple worked with Khronos to build OpenCL as a general competitor. The industry didn't come together to support it though, and eventually major stakeholders abandoned it altogether. Those ~10 wasted years were spent on Nvidia's side refining their software offerings and redesigning their GPU architecture to prioritize AI performance over raster optimization. Meanwhile Apple and AMD were pulling the rope in the opposite direction, trying to optimize raster performance at all costs. This means that Nvidia is selling a relatively unique architecture with a fully-developed SDK, industry buy-in and relevant market demand. Getting AMD up to the same spot would force them to reevaluate their priorities and demand a clean-slate architecture to-boot.
- DeathArrow 1y ago>In 2024, Python became the most popular programming language in the world — overtaking JavaScript — according to GitHub’s 2024 open source survey. I wonder why Python take over the world? Of course, it's easy to learn, it might be easy to read and understand. But it also has a few downsides: low performance, single threaded, lack of static typing.
- nhumrich 1y agoPerhaps performance, multi threading, and static typing are not the #1 things that make a language great. My guess: it's the community.
- chupasaurus 1y agoAll 3 are achieved in Python with a simple import ctypes /sarcasm
- timschmidt 1y agoUniversities seem to have settled on it for CSE 101 courses in the post-Java academic programming era.
- PeterStuer 1y agoIt's the ecosystem, specifically the huge amount of packages available for everything under the sun.
- diggan 1y ago> I wonder why Python take over the world? Not sure what "most popular programming language in the world" even means, in terms of existing projects? In terms of developers who consider it their main language? In terms of existing actually active projects? According to new projects created on GitHub that are also public? My guess is that it's the last one, which probably isn't what one would expect when hearing "the most popular language in the world", so worth keeping in mind. But considering that AI/ML is the hype today, and everyone want to get their piece of the pie, it makes sense that there is more public Python projects created on GitHub today compared to other languages, as most AI/ML is Python.
- CapsAdmin 1y agoSlightly related, I had a go at doing llama 3 inference in luajit using cuda as one compute backend for just doing matrix multiplication https://github.com/CapsAdmin/luajit-llama3/blob/main/compute/gpu_cuda.lua https://github.com/CapsAdmin/luajit-llama3/blob/main/compute... While obviously not complete, it was less than I thought was needed. It was a bit annoying trying to figure out which version of the function (_v2 suffix) I have to use for which driver I was running. Also sometimes a bit annoying is the stateful nature of the api. Very similar to opengl. Hard to debug at times as to why something refuse to compile.
- andrewmcwatters 1y agoNeat, thanks for sharing!
- btown 1y agoThe GTC 2025 announcement session that's mentioned in this article has video here: https://www.nvidia.com/en-us/on-demand/session/gtc25-s72383/ https://www.nvidia.com/en-us/on-demand/session/gtc25-s72383/ It's a holistic approach to all levels of the stack, from high-level frameworks to low-level bindings, some of which is highlighting existing libraries, and some of which are completely newly announced. One of the big things seems to be a brand new Tile IR, at the level of PTX and supported with a driver level JIT compiler, and designed for Python-first semantics via a new cuTile library. https://x.com/JokerEph/status/1902758983116657112 https://x.com/JokerEph/status/1902758983116657112 (without login: https://xcancel.com/JokerEph/status/1902758983116657112 https://xcancel.com/JokerEph/status/1902758983116657112 ) Example of proposed syntax: https://pbs.twimg.com/media/GmWqYiXa8AAdrl3?format=jpg&name=large https://pbs.twimg.com/media/GmWqYiXa8AAdrl3?format=jpg&name=... Really exciting stuff, though with the new IR it further widens the gap that projects like https://github.com/vosen/ZLUDA https://github.com/vosen/ZLUDA and AMD's own tooling are trying to bridge. But vendor lock-in isn't something we can complain about when it arises from the vendor continuing to push the boundaries of developer experience.
- skavi 1y agoi’m curious what advantage is derived from this existing independently of the PTX stack? i.e. why doesn’t cuTile produce PTX via a bundled compiler like Triton or (iirc) Warp? Even if there is some impedance mismatch, could PTX itself not have been updated?
- cavisne 1y agoIn the presentation they said eventually kernels can share SIMT (PTX) and TileIR but not at launch. It seems pretty mysterious why they don't just emit PTX, I would guess they are either taking the opportunity to clean things up for ML tensorcore workloads or there is some HW specific features coming that they only want to enable through TileIR.
- skavi 1y agoif i were to lean into cynicism, i might suggest this choice was meant to increase the effort required to reimplement cuda for other cards.
- lunarboy 1y agoDoesn't this mean AMD could make the same python API that targets their hardware, and now Nvidia GPUs aren't as sticky?
- KeplerBoy 1y agoNothing really changed. AMD already has a c++ (HIP) dialect very similar to CUDA, even with some automatic porting efforts (hipify). AMD is held back by the combination of a lot of things. They have a counterpart to almost everything that exists on the other side. The things on the AMD side are just less mature with worse documentation and not as easily testable on consumer hardware.
- pjmlp 1y agoThey could, AMD's problem is that they keep failing on delivery.
- jmward01 1y agoThis will probably lead to what, I think, python has led to in general: A lot more things tried quicker and targeted things that stay in a faster language. All in all this is a great move. I am looking forward to playing with it for sure.
- chrisrodrigue 1y agoPython is really shaping up to be the lingua franca of programming languages. Its adoption is soaring in this FOSS renaissance and I think it's the closest thing to a golden hammer that we've ever had. The PEP model is a good vehicle for self-improvement and standardization. Packaging and deployment will soon be solved problems thanks to projects such as uv and BeeWare, and I'm confident that we're going to see continued performance improvements year over year.
- silisili 1y ago> Packaging and deployment will soon be solved problems I really hope you're right. I love Python as a language, but for any sufficiently large project, those items become an absolute nightmare without something like Docker. And even with, there seems to be multiple ways people solve it. I wish they'd put something in at the language level or bless an 'official' one. Go has spoiled me there.
- horsawlarway 1y agoHonestly, I'm still incredibly shocked at just how bad Python is on this front. I'm plenty familiar with packaging solutions that are painful to work with, but the state of python was shocking when I hopped back in because of the available ML tooling. UV seems to be at least somewhat better, but damn - watching pip literally download 20+ 800MB torch wheels over and over trying to resolve deps only to waste 25GB of bandwidth and finally completely fail after taking nearly an hour was absolutely staggering.
- SJC_Hacker 1y agoPython was not taken seriously as something you actually shipped to non-devs. The solution was normally "install the correct version of Python on the host system". In the Linux world, this could be handled through Docker, pyenv. For Windows users, this meant installing a several GB distro and hoping it didn't conflict with what was already on the system.
- whycome 1y ago
- ryao 1y agoCUDA was born from C and C++ It would be nice if they actually implemented a C variant of CUDA instead of extending C++ and calling it CUDA C.
- swyx 1y agowhy is that impt to you? just trying to understand the problem you couldnt solve without a C-like
- kevmo314 1y agoA strict C variant would indeed be quite nice. I've wanted to write CUDA kernels in Go apps before so the Go app can handle the concurrency on the CPU side. Right now, I have to write a C wrapper and more often than not, I end up writing more code in C++ instead. But then I end up finding myself juggling mutexes and wishing I had some newer language features.
- ryao 1y agoI want to write C code, not C++ code. Even if I try to write C style C++, it is more verbose and less readable, because of various C++isms. For example, having to specify extern “C” to get sane ABI names for the Nvidia CUDA driver API: https://docs.nvidia.com/cuda/cuda-driver-api/index.html https://docs.nvidia.com/cuda/cuda-driver-api/index.html Not to mention that C++ does not support neat features like variable sized arrays on the stack.
- pjmlp 1y ago
- thegabriele 1y agoCould pandas benefit from this integration?
- binarymax 1y agoPandas uses numpy which uses C. So if numpy used CUDA then it would benefit.
- ashvardanian 1y agoCheck out CuDF. I've mentioned them in another comment on this thread.
- thegabriele 1y agoIt seems already a mature and seamless solution. Thank you
- system2 1y agoWhat do you have in mind that Pandas will benefit from cuda cores?
- thegabriele 1y agoI'm absolutely not a pro coder, i just crunch numbers and i would love to see some improvements speed-wise in dataframe operations.
- math_dandy 1y agoIs NVIDIA's JIT-based approach here similar JAX's, except targeting CUDA directly rather than XLA? Would like to know how these different JIT compilers relate to one another.
- WhereIsTheTruth 1y agopython is the winner, turning pseudo code into interesting stuff it's only the beginning, there is no need to create new programming languages anymore
- system2 1y agoI heard the same thing about Ruby, Go, TypeScript, and Rust. (Even JS at one point when NodeJS was super popular a few years ago). There will be new shiny things, but of course, my choice is Python too.
- steelbrain 1y agoSee also: https://tinygrad.org/ https://tinygrad.org/ Reverse-engineered python-only GPU API, works with not only CUDA but Also AMD's ROCm Other runtimes: https://docs.tinygrad.org/runtime/#runtimes https://docs.tinygrad.org/runtime/#runtimes
- odo1242 1y agoTechnically speaking, all of this exists (including the existing library integration and whatnot) through CuPy and Numba already, but the fact that it’s getting official support is cool.
- matt3210 1y agoWe should find a new word since GPU is is from back when it was used for graphics.
- martinsnow 1y agoDisagree. It's the name of the component and everyone working with it knows what capabilities it has. Perhaps you got old.
- SJC_Hacker 1y agoGeneral Processing Unit Greater Processing Unit Giant Processing Unit
- no_wizard 1y agoCouple more thoughts: Galloping Processing Unit Grape Processing Unit Gorge Processing Unit Gaggle Processing Unit Grand Processing Unit Giraffe Processing Unit Gaping Processing Unit
- chrisrodrigue 1y agoLinear Algebra Unit?
- ashvardanian 1y agoCuTile, in many ways, feels like a successor to OpenAI's Triton... And not only are we getting tile/block-level primitives and TileIR, but also a proper SIMT programming model in CuPy, which I don't think enough people noticed even at this year's GTC. Very cool stuff! That said, there were almost no announcements or talks related to CPUs, despite the Grace CPUs being announced quite some time ago. It doesn't feel like we're going to see generalizable abstractions that work seamlessly across Nvidia CPUs and GPUs anytime soon. For someone working on parallel algorithms daily, this is an issue: debugging with NSight and CUDA-GDB still isn't the same as raw GDB, and it's much easier to design algorithms on CPUs first and then port them to GPUs. Of all the teams in the compiler space, Modular seems to be among the few that aren't entirely consumed by the LLM craze, actively building abstractions and languages spanning multiple platforms. Given the landscape, that's increasingly valuable. I'd love to see more people experimenting with Mojo — perhaps it can finally bridge the CPU-GPU gap that many of us face daily!
- saagarjha 1y agoI mean, it doesn’t really make sense to unify them. CPUs and GPUs have very different performance characteristics and you design for them differently depending on what they let you do. There’s obviously a common ground where you can design mostly good interfaces to do things ok (I’ll argue PyTorch is that) but it’s not really reasonable to write an algorithm that is hobbled on CPUs for no reason because it assumes that synchronizing between execution contexts is super expensive.
- jms55 1y ago> And not only are we getting tile/block-level primitives and TileIR As someone working on graphics programming, it always frustrates me to see so much investment in GPU APIs _for AI_, but almost nothing for GPU APIs for rendering. Block level primitives would be great for graphics! PyTorch-like JIT kernels programmed from the CPU would be great for graphics! ...But there's no money to be made, so no one works on it. And for some reason, GPU APIs for AI are treated like an entirely separate thing, rather than having one API used for AI and rendering.
- crazygringo 1y agoVery curious how this compares to JAX [1]. JAX lets you write Python code that executes on Nvidia, but also GPUs of other brands (support varies). It similarly has drop-in replacements for NumPy functions. This only supports Nvidia. But can it do things JAX can't? It is easier to use? Is it less fixed-size-array-oriented? Is it worth locking yourself into one brand of GPU? [1] https://github.com/jax-ml/jax https://github.com/jax-ml/jax
- odo1242 1y agoWell, the idea is that you’d be writing low level CUDA kernels that implement operations not already implemented by JAX/CUDA and integrate them into existing projects. Numba[1] is probably the closest thing I can think of that currently exists. (In fact, looking at it right now, it seems this effort from Nvidia is actually based on Numba) [1]: https://numba.readthedocs.io/en/stable/cuda/overview.html https://numba.readthedocs.io/en/stable/cuda/overview.html
- rahimnathwani 1y agoHere is the repo: https://github.com/NVIDIA/cuda-python https://github.com/NVIDIA/cuda-python
- melodyogonna 1y agoWill be interesting to compare the API to what Modular has [1] with Mojo. 1. https://docs.modular.com/mojo/stdlib/gpu/ https://docs.modular.com/mojo/stdlib/gpu/
- hingusdingus 1y agoHmm wonder what vulnerabilities will now be available with this addition.
- soderfoo 1y agoKind of late to the game, but can anyone recommend a good primer on GPU programming?
- no_wizard 1y agoWhat makes Python such a target for these kind of things? I've noticed alot of projects add Python support like this. Does the Python codebase allow for it to compile down to different targets easier than others?
- saagarjha 1y agoThere’s a lot of existing Python code in this space and many ML researchers are comfortable in Python.
- jdeaton 1y agoPytorch???
- exesiv 1y ago[flagged]