3 ms·
I’d start with CUDA, because knowing what a chip does won’t click until you see how it can be programmed to do massive parallel computation and matmul. I read
by binarymax 3y ago
I’d start with CUDA, because knowing what a chip does won’t click until you see how it can be programmed to do massive parallel computation and matmul.
I read the first book in this list about 10 years ago, and though it’s pretty old the concepts are solid.
https://developer.nvidia.com/cuda-books-archive https://developer.nvidia.com/cuda-books-archive
- dogma1138 3y agoCUDA abstracts most of the parallelism, the magic of CUDA is it gave developers a C/C++ API or language if you will that doesn’t really requires them to think about that they can continue writing their problems as they did when programming for mostly single core single threaded CPUs back in the day and CUDA takes care of the rest. Even “manual” CUDA optimizations deal more with concurrency and data residency than parallelism and even those are usually limited to following the compute guide for your specific hardware and feature set and the driver does the majority of the heavy lifting.