5 ms·
After having attended a few CUDA workshops at NVIDA’s latest GTC, I was inspired to continue learning CUDA on my own. To do so I decided to build John Conway’s
by brendanrayw 5y ago
After having attended a few CUDA workshops at NVIDA’s latest GTC, I was inspired to continue learning CUDA on my own. To do so I decided to build John Conway’s famous “Game of Life” and use CUDA to accelerate the program. I explore multiple different CUDA techniques including managed memory, pinned memory, multiple streams, and asynchronous memory transfers.
- gtn42 5y agoNice, thanks for sharing your experience!
- IdiocyInAction 5y agoNice. I took a CUDA course at uni where I built a neural network and a physics simulation. Optimizing them was very involved, but ultimately very cool; I learned a ton of stuff. I'd love to work with CUDA in practice, but there's not that many jobs around.
- rrss 5y agothis was a fun read, thanks for sharing. FYI, the transfers from pageable memory almost certainly do not go to the storage device in your system, unless you have high memory pressure. "pageable" (as a cuda-ism) does mean that the buffer may be paged out to storage, but as a result it means that (more importantly) even if the buffer is in RAM, the GPU cannot access it directly. so for pageable copies the flow is probably not: storage → buffer in RAM → device, but rather: original buffer in RAM (inaccessible to the device) → intermediate buffer in RAM (accessible to the device) → device. also, in several places you use the term 'stack' where I think it should just be 'RAM' / main memory.
- brendanrayw 5y agoThanks for reading! I appreciate the feedback and the info, I'll keep that in mind.
- joe_the_user 5y agoThanks for your effort! I really like the idea, it's similar to a more ambitious project I'm thinking of. And I do have questions Is your board a giant two-dimensional array in memory? Are your threads/kernels reading from this array and then writing back to it? Do you do synchronization to make sure reads happen before the later rights? Do you do any verification that your transition happen correctly? Do you have an estimate for time spend in - transfer from global GPU memory to each kernel, calculations in the kernel, and time spent idling through synchronization (assuming you do it).