9 ms·
Introduction to CUDA programming for Python developers
- spps11 2y agoThanks for sharing, enjoyed reading it! I have a slightly tangential question: Do you have any insights into what exactly DeepSeek did by bypassing CUDA that made their run more efficient? I always found it surprising that a core library like Cuda, developed over such a long time, still had room for improvement—especially to the extent that a seemingly new team of developers could bridge the gap on their own.
- t55 2y agoThey basically ditched CUDA and went straight to writing in PTX, which is like GPU assembly, letting them repurposing some cores for communication to squeeze out extra performance. I believe that with better AI models and tools like Cursor, we will move to a world where you can mold code ever more specific to your use case to make it more performant.
- spps11 2y agogot it, thanks for explaining. > with better AI models and tools like Cursor, we will move to a world where you can mold code ever more specific to your use case to make it more performant what do you think the value of having the right abstraction will be in such a world?
- t55 2y agoI think that for at least for us dumb humans with limited memory, having good abstractions makes things much easier to understand
- spps11 2y agoYes, but I wonder how much of this trait is carried over to the LLMs from us.
- t55 2y agowhat do you mean, the LLM abstracting things for us while we speak to it?
- spps11 2y agoNo I meant something else. As you said: us humans love clean abstractions. We love building on top of them. Now LLMs are trained on data produced by us. So I wonder if they would also inherit this trait from us and end up loving good abstractions, and would find it easier to build on top of them. Other possibility is that they end up move-37ing the whole abstraction shebang. And find that always building something up bespoke, from low-level is better than constraining oneself to some general purpose abstraction.
- tomnipotent 2y agoIt's an interesting idea. If code is ever updated by an LLM, does it benefit from using abstractions? After all they're really a tool for us lowly sapients to aid in breaking down complex problems. Maybe LLM's will create their own class of abstractions, diverse from our own but useful for their task.
- t55 2y agoah gotcha. I think that with the new trend of RLing models, the move 37 may come up sooner than we think -- just provide the pretrained models some outcome-goal and the way it gets there may use low-level code without clean abstractions
- suresk 2y agoAre you sure they ditched CUDA? I keep hearing this, but it seems odd because that would be a ton of extra work to entirely ditch it vs selectively employing some ptx in CUDA kernels which is fairly straightforward. Their paper [1] only mentions using PTX in a few areas to optimize data transfer operations so they don't blow up the L2 cache. This makes intuitive sense to me, since the main limitation of the H800 vs H100 is reduced nvlink bandwidth, which would necessitate doing stuff like this that may not be a common thing for others who have access to H100s. 1. https://arxiv.org/abs/2412.19437 https://arxiv.org/abs/2412.19437
- t55 2y agoI should have been more precise, sorry. Didn't want to imply they entirely ditched CUDA but basically circumvented it in a few areas like you said.
- pjmlp 2y agoTargeting directly PTX is perfectly regular CUDA, and used by many toolchains that target the ecosystem. CUDA is not only C++, as many mistake it for.
- saagarjha 2y agoThey didn’t. They used PTX, which is what CUDA C++ compiles down to, but which is part of the CUDA toolchain. All major players have needed to do this because the intrinsics for the latest accelerators are not actually exposed in the C++ API, which means using them requires inline PTX at the very minimum.
- t55 2y agoRelated: https://sakana.ai/ai-cuda-engineer/ https://sakana.ai/ai-cuda-engineer/ https://www.reddit.com/r/MachineLearning/comments/1itqrgl/p_sakana_ai_released_cuda_ai_engineer/ https://www.reddit.com/r/MachineLearning/comments/1itqrgl/p_...
- saagarjha 2y agoWasn’t this a bunch of kernels that didn’t work?
- tsunego 2y ago[dead]
- t55 2y agoWhat do you mean?
- pavelstoev 2y agoThe hallucinated code was reusing memory buffers filled with previous results so not performing the actual computations. When this was fixed the AI generated code was like 0.3x of the baseline.
- neodypsis 2y agoIt is mentioned on section "Limitations and Bloopers" of the page [0]: > Combining evolutionary optimization with LLMs is powerful but can also find ways to trick the verification sandbox. We are fortunate to have Twitter user @main_horse help test our CUDA kernels, to identify that The AI CUDA Engineer had found a way to “cheat”. The system had found a memory exploit in the evaluation code which, in a small percentage of cases, allowed it to avoid checking for correctness (...) 0. https://sakana.ai/ai-cuda-engineer https://sakana.ai/ai-cuda-engineer
- rnrn 2y ago
- ferguess_k 2y agoStupid question: Is there any chance that I, as an engineer, can get away from learning the Math side of AI but still drill deeper into the lower level of CUDA or even GPU architecture? If so, how do I start? I guess I should learn about optimization and why we chose to use GPU for certain computations. Parallel question: I work as a Data Engineer and always wonder if it's possible to get into MLE or AI Data Engineering without knowing AI/ML. I thought I only need to know what the data looks like, but so far I see every job description of an MLE requires background in AI.
- mlazos 2y agoYou don’t need to be deep in designing NNs and the theory behind them, but I would say you should be able to take some linear algebra equations and be able to map them to the GPU arch. This does require some knowledge of the math being used. Luckily it’s mostly high-school/college level math. Starting with the CUDA and tritonlang docs are a good starting point for an introduction. They’ll teach you about common optimizations like tiling, thread swizzling and maximizing cache utilization.
- danielmarkbruce 2y agoYes. They are largely unrelated. Just go to Nvidia's site and find the docs. Or there are several books (look at amazon). A "background in AI" is a bit silly in most cases these days. Everyone is basically talking about LLMs or multimodal models which in practice haven't been around long. Sebastian Raschka has a good book about building an LLM from scratch, Simon Prince has a good book on deep learning, Chip Huyen has a good book on "AI engineering". Make a few toys. There you have a "background". Now if you want to really move the needle... get really strong at all of it, including PTX (nvidia gpu assembly, sort of). Then you can blow people away like the deep seek people did...
- t55 2y agoAgreed, Rashka's book is amazing and will probably become the seminal book on LLMs
- 2y ago
- lukaspetersson 2y agoI needed this
- t55 2y agoHehe glad you did!
- rtkal10 2y agoInterestingly, the CUDA implementations are more readable than the pytorch ones.
- m_kos 2y agoSince this is on PySpur's website, does anyone have experience with these UI tools for AI agents like PySpur and n8n? I am looking for something to help me prototype a few ideas for fun. I would have to self-host it ($), so I would prefer something relatively easy to configure like Open Hands.
- spps11 2y agopyspur is apache 2. it is free to self-host.
- t55 2y agoDisclaimer: I work on pyspur I'd recommend pyspur if you seek 1) More AI-native features eg. Evals, RAG, or even UI decisions like seeing outputs directly on the canvas when running on the agent 2) Truly open-source Apache license 3) Python-based (in the sense that you can run and extend it via python) On the other hand, n8n is 1) more mature for traditional workflows 2) offering overall more integrations (probably every single integration you can think of) 3) TypeScript based and runs on Node.js
- m_kos 2y agoThanks for replying. Do you know when your docs will be a bit more comprehensive? Right now, there is very little information and some links don't work, e.g., Next Steps on this page: https://docs.pyspur.dev/quickstart https://docs.pyspur.dev/quickstart
- t55 2y ago> Do you know when your docs will be a bit more comprehensive? Yes, we're actively working on this, and we should have some more pages by next week. If you have any questions, you can always shoot us an email: founders@pyspur.dev or join our Discord. > some links don't work, e.g., Next Steps on this page This might be confusing, the cards below "After installation, you can:" are not meant to be links. Thanks for making us aware, we will improve the wording.
- nitrogen99 2y agoIf you are a Python dev, why not just use Triton?
- saagarjha 2y agoTriton is somewhat limited in what it supports, and it’s not really Python either.
- t55 2y agoTriton sits between CUDA and PyTorch and is built to work smoothly within the PyTorch ecosystem. In CUDA, on the other hand, you can directly manipulate warp-level primitives and fine-tune memory prefetching to reduce latency in eg. attention algorithms, a level of control that Triton and PyTorch don't offer AFAIK.
- pjmlp 2y agoMLIR extensions for Python do though, as far as I could tell from LLVM developer meeting.
- 6gvONxR4sf7o 2y agoMLIR is one of those things everyone seems to use, but nobody seems to want to write solid introductory docs for :( I've been curious for a few years now to get into MLIR, but I don't know compilers or LLVM, and all the docs I've found seem to assume knowledge of one or the other. (yes this is a plea for someone to write an 'intro to compilers' using MLIR)
- pjmlp 2y agoNot sure if you will be able to follow along, but here it is what I was talking about, "PyDSL: A MLIR DSL for Python developers" https://www.youtube.com/watch?v=iYLxgTRe8TU https://www.youtube.com/watch?v=iYLxgTRe8TU "PyDSL, a subset of Python for constructing affine & transform dialects" https://www.youtube.com/watch?v=nmtHeRkl850 https://www.youtube.com/watch?v=nmtHeRkl850 And MLIR channel, https://www.youtube.com/@MLIRCompiler https://www.youtube.com/@MLIRCompiler
- tsunego 2y ago[dead]
- LegNeato 2y agoAlso check out https://github.com/rust-gpu/rust-gpu https://github.com/rust-gpu/rust-gpu and https://github.com/rust-gpu/rust-cuda https://github.com/rust-gpu/rust-cuda
- t55 2y agothis looks really cool and i love rust. just a matter of time until everything runs on rust.
- the__alchemist 2y agoRust-Cuda is broken and has been for years.`cudarc` is the [only?] working one.
- LegNeato 2y agoI am in the process of rebooting it: https://rust-gpu.github.io/blog/2025/01/27/rust-cuda-reboot/ https://rust-gpu.github.io/blog/2025/01/27/rust-cuda-reboot/
- musicale 2y agoWhat Jensen giveth, Guido taketh away.
- t55 2y agolol. i guess this tutorial is about cutting out guido ;)
- signa11 2y agothis book: Programming Massively Parallel Processors by Wen-mei W. Hwu , David B. Kirk , Izzat El Hajj seems to be tailor mode for folks transitioning from cpu -> gpu arch.
- t55 2y agoYes, it is great for key concepts but a bit outdated. Hence we added an LLM/FA section in the linked post!
- whatever1 2y agoAny idea what changed recently and we can have end to end simulations (with branches) in the gpu (eg isaac sim) vs in the past where simulations were a cpu thing ?
- jamiejquinn 2y agoAlways been possible, but now the time cost of moving data between the GPU and CPU memory is too high to ignore. Branching may be slower on the GPU but it's still faster than moving data to the CPU for a time then back. The maturation of direct GPU-GPU transfers over the network also helped enable GPU-only MPI codes.
- android521 2y agopyspur graph is cool, is there a startup building this kind of product but in typescript?
- AchintyaAshok 2y agoThanks for unraveling this!
- t55 2y agoyou're welcome!
- ultrasounder 2y agoVery nice-write up. The in-line quiz, which i think is AI generated(QnA) is very useful to test understanding. Wish all tutorials incorporated that feature.
- t55 2y agothank you!
- ralphc 2y agoAre all the CUDA tutorials geared towards AI or are there some, for example, like regular scientific computing? Airflow over wings and things that you used to see for high-performance computing would be fun to try.