3 ms·
Actually, initial CUDA trials can be quite disappointing, unless heavy effort is spent in understanding the architecture, so to be able to exploit it. Then you
by Create 17y ago
Actually, initial CUDA trials can be quite disappointing, unless heavy effort is spent in understanding the architecture, so to be able to exploit it. Then you are back at square one: (almost?) a rewrite. GSI has nice examples about their studies on different architectures. Then again, the real danger when going with non mass-consumer items, is that the product lifecycle ends before you finish your project. A weird but teaching example, even in case of mass produced technology: a high-performance switch maker used a chip that was built for a mainstream game console. As the console switched generations, the part availability dried up...
I just hope that OpenCL takes off and becomes really open (read free as in free drivers and code). Once familiarized with the quirks of a new architecture (read non x86-alike) then it would require less investment to draw computing power (btw. Nvidia is a power-hog, significantly reducing its appeal in large data centres).
- profquail 17y agoI've worked with CUDA a pretty good amount, so I feel obliged to point out that most of initial effort put into porting an algorithm to CUDA is not spent learning the architecture, but parallelizing the algorithm (which is something you would have to deal with even if just porting your code to multi-core CPU). Once you have figured out how to parallelize the algorithm, then you need to worry about the architecture, but only if you need to eek out every last bit of performance. In most viable cases, you should be able to get a pretty decent speedup without tearing your hair out. In any case, OpenCL is based on CUDA (the driver API), and they are practically the same (compare the reference manuals and you'll see what I mean). Also, GPU computing is much, much more efficient than CPU computing (in terms of FLOPS/Watt). A high-end GPU (say, a Tesla C1060) can probably pull between 200-300 watts (~2x what a high-end Xeon uses), but can do over 1TFLOP with sufficiently optimized code.