4 ms·
This is where the "100x myth" comes from. If you write a naive scalar algorithm in C, and then you write (almost) the same thing in CUDA, you can get a ~100x sp
by dsharlet 10y ago
This is where the "100x myth" comes from. If you write a naive scalar algorithm in C, and then you write (almost) the same thing in CUDA, you can get a ~100x speedup. The problem is that most of that speedup is coming from the lack of parallelism and vectorization in the C implementation, which CUDA gives you "for free".
The thing is, this is still a real win for CUDA, because most people aren't going to write highly optimized C. It's just not a hardware win, it's a compiler win. The thing that most people get out of CUDA is not an advantage from running on a GPU, but an advantage of a smart compiler/language design. Most of the same people would get all but 1-5x of the benefits from using something like ISPC[1], which gives a similar programming model as CUDA implemented on the CPU.
In my experience (from the Tesla/GTX 200 era, though I doubt things have changed that much now), such a small boost in performance is not worth the hassle of transferring data to/from the GPU, the (lack of) virtualization, and driver shenanigans/support issues (at one point, I had to suggest someone buy a fake monitor dongle to plug into his GPU to be able to use my code...).
Anyways, it's good to see this thread, people these days think you are crazy for not using the GPU to do big compute workloads.
1. https://ispc.github.io/ https://ispc.github.io/
- hughperkins 10y agohave you actually tried writing highly optimizer cuda? its really hard. nothing is "for free". a lot of the caching (not all) has to be hand coded. you can see how much attention to detail scott gray puts into his kernels eg at https://github.com/NervanaSystems/maxas https://github.com/NervanaSystems/maxas
- jacquesm 10y agoNot that much harder than writing hand optimized assembly. There are a few more quirks to keep in mind and you need to have the memory lay-out and access patterns down otherwise it will not give you the boost you expect. In a way I really love CUDA for this reason, it allows me to re-cycle all my old optimization skills and do something useful with them.
- glangdale 10y agoCUDA is pretty neat; I worked with an early beta (0.7?) and have tracked it on and off over time. It's not magic, though. The point is that GPGPU is really great for some things and pretty ordinary for others. 'Embarrassingly parallel' algorithms do well, as do things that are heavy consumers of memory bandwidth. Like you say, anyone who needs frequent transfers to/from CPU may not do as well. Highly optimized C and CUDA are, to me, the same concept - to get the best results, you have to understand the underlying machine to get the best results. While your resources are different, it's a very similar skill - to codesign an algorithm with the deep knowledge of how either a modern CPU or a GPGPU actually works. I'm no longer current as a GPGPU programmer but it wouldn't surprise me that the same 10-20x that an expert can get with astute design for modern architecture is available on GPGPU as well (i.e. the difference between a naive and a sophisticated design on the same platform). Personally I regard that kind of algorithmic codesign as one of the most fun things you can do with your clothes on, but that's probably just me.
- jacquesm 10y agoAh, but when you can leave most of your data on the GPU side and only transfer little bits to solve your problem GPUs absolutely shine. Some problems are well suited to this kind of architecture and some not, if you're in luck the gains can be significant and it is definitely something you should be aware of and weigh when you're writing something that is heavy on computation.