4 ms·
CUDA Books
- phoronixrly 5mo agoIn an age when your company mandates you to raise your productivity right now with hundreds of percentage points using LLMs, how do you find an excuse to sit down and read a book?
- q8zd3 5mo agoIt feels like a dirty secret, doesn't it?
- phoronixrly 5mo agoYeah, corps don't want you to know how to code, they want you to be a prompter...
- fileeditview 5mo agoDon't you read while your agents are doing all the work for you? /s
- hartator 5mo agoOr make your agents do the reading for you!
- deleted 5mo ago[deleted]
- mohamedkoubaa 5mo agoAnthropunk
- signa11 5mo agonot on company time ?
- pjmlp 5mo agoAs always, on private time, if available, otherwise wait when LLM connection breaks down.
- chrsw 5mo ago"AI Systems Performance Engineering" might deserve a mention, even though it's not strictly CUDA.
- zparky 5mo agoI liked going through https://www.olcf.ornl.gov/cuda-training-series/ https://www.olcf.ornl.gov/cuda-training-series/ for an intro and some fundamentals.
- lacedeconstruct 5mo agoGoing through books after this one was a breeze
- juvoly 5mo agoIncreasingly (for instance ADSP podcast [1]) those in nvidia's inner circle are advocating against writing your own CUDA kernels. (Unless that's your full time job at nvidia, that is). [1] https://adspthepodcast.com/2024/08/30/Episode-197.html https://adspthepodcast.com/2024/08/30/Episode-197.html
- dahart 5mo agoIt’s not about whether you work at Nvidia. Avoid writing CUDA kernels if there are higher level libraries that do what you need. Do write CUDA kernels if you want to learn how, or if you need the low level control, or to micro-optimize. Being able to fuse kernels to avoid memory traffic or get better specialization is also a reason to reach for raw CUDA. Just consider what’s the right tool for the job…
- saagarjha 5mo agoI don't think writing CUDA is a good way to do this tbh
- nnevatie 5mo agoTo do what? If you need the highest performance GPU kernel performance on NVidia HW, using CUDA is the way to go.
- saagarjha 5mo agoWriting efficient CUDA code is very, very difficult; most CUDA code is not actually good at utilizing the hardware. It is much easier to write performant code in higher level languages (and most people are doing exactly this).
- dahart 5mo agoThat all depends on what you’re doing. Like I said, if a high level lang or lib supports and fits your goal well, then yes you should use it. I don’t know what most people are doing, but it’s fair to say that a lot of people can use a higher level language. If you’re trying to learn CUDA, then using a higher level language is not the best approach. If you already used a high level language and found that your performance is lacking and could be better if you could fuse some of your kernels, and avoid some of the memory round-trips, then moving to something lower level is called for. I’m suggesting it’s better to think about your goals for one minute and understand the basic choices than it is to assume there’s something that works for everyone’s goals, and higher level languages don’t meet everyone’s goals.
- pwython 5mo agoFirst one I clicked on is 404: Programming Massively Parallel Processors: A Hands-on Approach (3rd Edition) https://www.cambridge.org/core/books/programming-in-parallel-with-cuda/9781108855273 https://www.cambridge.org/core/books/programming-in-parallel...
- synergy20 5mo agothe newest is 4th ed i think
- cdavid 5mo agoA fifth edition has been out recently: https://shop.elsevier.com/books/programming-massively-parallel-processors/hwu/978-0-443-43900-1 https://shop.elsevier.com/books/programming-massively-parall... I started learning about GPU and CUDA from this book recently, and I agree the writing is confusing, and code examples have errors. However, it is still a nice reference about many types of algorithms for heterogeneous memory devices, it helped me understand better some patterns for CPUs.
- dahart 5mo agoRegarding the section on Python and high-level CUDA, anyone interested should maybe first take a peek at Warp, which I’m guessing is too new to have a book yet. Warp lets you write CUDA kernels directly in Python, and it’s a breeze to get started. https://github.com/nvidia/warp https://github.com/nvidia/warp
- tirutiru 5mo agoIt's a bit confusing now with Numba Cuda also being officially maintained by Nvidia. Also Cuda Python, which looks older. Which of these - warp, numba, cp, is the best bet for a beginner? https://nvidia.github.io/numba-cuda/ https://nvidia.github.io/numba-cuda/ https://developer.nvidia.com/cuda/python https://developer.nvidia.com/cuda/python
- dahart 5mo agoI haven’t tried them all, but I suspect Warp is the easiest; it’s ridiculously easy. I’m sure there are some tradeoffs, so once you learn a little CUDA in Python it might make sense to switch from Warp to Numba or CP depending on what you’re doing.
- dandanua 5mo agoYou can also write CUDA kernels directly in Julia using CUDA.jl. I basically learned CUDA programming by experimenting in Julia with the help of LLMs.
- somethingsome 5mo agoHaving read or at least skimmed most of those books, I think the best intro is 'CUDA Programming: A Developer's Guide to Parallel Computing with GPUs' Massively Parallel Processors: A Hands-on Approach is not really good in my opinion, many small mistakes and confusing sentences (even when you know cuda). CUDA by Example: An Introduction to General-Purpose GPU Programming is too simple and abstract too much the architecture. Next year I'm planning to start writing a cuda book that starts by engineering the hardware, and goes up to the optimization part on that harware (which is basically a nvidia card) including all the main algorithms (except for graphs). I'm already teaching the course in this way at uni, and it is quite successful among students.
- synergy20 5mo agothe first book was published in 2012,is it too outdated?
- somethingsome 5mo agoNot really, Hardware didn't really change that much, of course you'll not find Tensor or raytracing cores, but you will have a very solid grasp of gpu programming and the cuda language (that didn't change that much either), and then you can easily learn those more modern things with blog posts or even, at worst, chatgpt.
- jpgvm 5mo agoYeah pretty much this. I would separate the knowledge into maybe 3 distinct buckets. The baseline: device/host boundary, SIMT programming etc. The intermediate: kernel architecture, CUDA graph vs persistent kernels, warp specialisation/divergence avoidance techniques etc. The advanced: architecture specifics so tcgen05, TMA, SMEM/HBM, memory throughput vs compute biases in various arch impls., GEMM, FHMA, all the tricks that make modern fused kernels very fast. Also would bucket most GPU Direct RDMA/GPU NetIO/friends here too. The baseline hasn't changed much and probably won't, the intermediate knowledge has also remained pretty reliably stable for ~10 years with only things like graphs changing stuff. Tile might become more relevant than it is today but for now CUDA, cuBLAS, friends are where it's worth investing knowledge.
- brcmthrowaway 5mo agoAny good MOOCs on Parallel programming/NVIDIA?
- fwx 5mo agohttps://www.youtube.com/playlist?list=PLzn6LN6WhlN06hIOA_ge6SrgdeSiuf9Tb https://www.youtube.com/playlist?list=PLzn6LN6WhlN06hIOA_ge6...
- xiongdinggua 5mo agohttps://ppc.cs.aalto.fi/ https://ppc.cs.aalto.fi/
- fwx 5mo agoDoes anyone know of any good resources for the newer paradigms like cuTile?
- SkiFreeWin3 5mo agoI wish the README had a solid “what cool things you can do with this” right at the top. In this day and age when programming is so accessible, why not have a more tempting pitch than just book titles categorized by difficulty.
- fransje26 5mo agoI'll give you the TL;DR: With CUDA, you can make Nvidia GPUs go brrrr. Oh. And thereby, incidentally conquer the compute world.
- saagarjha 5mo agoProbably worth noting that writing performant kernels for modern Nvidia hardware looks almost nothing like what the books from 2012 are going to teach you. You can read them for fun if you'd like but they're basically irrelevant.
- wces 5mo agoThis is highly condensed video of all important concepts in CUDA from Stephen Jones, one of the CUDA architects: https://www.youtube.com/watch?v=QQceTDjA4f4 https://www.youtube.com/watch?v=QQceTDjA4f4 Understand everything he talks about and you understand CUDA.
- qzgrid37 5mo ago[dead]
- cold_harbor 5mo agofor LLM work, reading the Flash Attention and vLLM kernel source taught me more than any book. real code makes memory hierarchy concrete — books stay too abstract.
- dandanua 5mo agoThe story of Flash Attention is the best manifestation of power and difficulty of GPU programming. This page gives a nice overview of it https://aiwiki.ai/wiki/flash_attention https://aiwiki.ai/wiki/flash_attention
- aaqaishtyaq 5mo agoI really need to buy nvidia GPUs to be able to learn CUDA.
- adrian_b 5mo agoFortunately, unlike with the AMD GPUs, using one of the cheaper NVIDIA GPUs is sufficient for learning CUDA, because CUDA works similarly on all models. An expensive NVIDIA GPU is required only if your purpose is not just to learn, but to actually do useful graphics or ML/AI work.