5 ms·
What makes you think that CUDA's value prop is memory management? I have seen no indication that implementing the memory allocation system is the hard part of a
by ColonelPhantom 2y ago
What makes you think that CUDA's value prop is memory management? I have seen no indication that implementing the memory allocation system is the hard part of any GPGPU system, and HIP is mostly source-compatible with CUDA, so it's not like the user has to do anything different memory management-wise to run on AMD.
- roenxi 2y agoI owned an AMD card. Everything high-level I wanted to do I could implement on the CPU fine, but I couldn't implement it on the GPU because it was too hard to operate on data in VRAM without crashing (been a while, I forget if it was the transfer or the operation that failed), and to get back good debug info because I needed to move data from VRAM to normal RAM. I was trying to learn at the time so efficiency wasn't really a factor, I just wanted things to work. They never did. CPU was easy, GPU was hard. The major difference I could spot though was all the time I was spending batting data around between different buffers, the APIs weren't really limiting me.
- kllrnohj 2y agoGetting good performance out of the GPU in general is regularly a memory problem. CUDA doesn't really help beyond documenting what is expected of you with regards to data layout and warp coalescing. https://docs.nvidia.com/cuda/cuda-c-programming-guide/#maximize-memory-throughput https://docs.nvidia.com/cuda/cuda-c-programming-guide/#maxim...
- roenxi 2y agoIf you have a similar document for how to do all that with AMD cards and could link it to me back in 2014 that'd be hugely appreciated. Because I made no headway at all on the subject. The library writers are obviously a lot better at all this than I am, but the experience of owning an AMD card was that they literally didn't and my experience left me believing that the reasons were more specific than "no CUDA". My experience was everything compute related caused crashes.
- saagarjha 2y agoCUDA doesn't help directly, but Nvidia will provide you with libraries to do what you want efficiently on their hardware so you don't have to.
- ColonelPhantom 2y agoOK, but isn't the memory management stuff part of the API? :) And with HIP, the user-facing memory model is virtually identical to CUDA to my knowledge. Converting a CUDA program to HIP, assuming no unsupported features or platform details are used, is basically just "do s/cu/hip/g". Also, HIP wasn't around in 2014. To program an AMD GPU back then, you were probably using OpenCL, which is nicely platform independent but also arguably somewhat archaic.
- paulmd 2y ago> and to get back good debug info because I needed to move data from VRAM to normal RAM. I was trying to learn at the time so efficiency wasn't really a factor I mean you're literally using the calvinball programming language - it's just a metafunctor in the language of macros, what's the big deal? literally my (sanitized) code from 10 years ago ```cuda __global__ void kernel_myKernel(int num_items, myTask_t * task_item, #if VALIDATION == 1 myDebug1_t * output_debug1_arr, myDebug2_t * output_debug2_arr, #endif myOutput_t * output_arr, int current_iteration) ``` those myDebug_t pointers are pointers to host memory, so transferring the data back to host memory is transparently managed by CUDA or Thrust, and you just have an #IF or #IFDEF block in the code that dumps intermediate state to the pointer if validation is enabled in the build. And then you can run whatever test suite, on the actual intermediate data that's happening inside your kernel, so you can test your invariants/assertions. Test-suite lite edition/just the part you need - but you're testing actual kernel invocations, not just theoretical. And you can go to the level of dumping intermediate state of prefix_sums/reductions (or generalized "intermediate work items") every warp/grid iteration if you want - it's debug mode. but granted this relies on the ability to have data dynamically piped between device and host memory spaces... which was a novel bit of syntactic sugar they added to the CUDA toolkit back in like 2012 (it was big around the Kepler era and it was a big feature on the Jetson TK1 too) so AMD might well have not had it. Or they might have said they had it but it just broke horribly if you used it. but this is literally like three lines of code, if you had just used NVIDIA instead of AMD, because the feature was there when you needed it. Technically can be accomplished by just allocating extra VRAM for the buffer and copying it back afterwards, but granted, more work there. What's the price of getting your work done? What's your time worth? That's always been the only moat of relevance that CUDA has had... and it's always been worth the expense.