5 ms·
GPU acceleration sounds all great.. until anyone actually generalizes the framework, implements applications, and measures even the most trivial workload. There
by nemonemo 13y ago
GPU acceleration sounds all great.. until anyone actually generalizes the framework, implements applications, and measures even the most trivial workload. There should be some niche applications you may get 2-10x benefit, but 1000x is either measurement error, apple-to-banana comparison, or a new marketing gimmick. Unless you have a very fitting application, it is very easy for GPU to have less performance than CPU. I would be very curious what application actually could have 1000x performance enhancement. Even memory bandwidth, the main benefit of GPU, does not differ by that amount.
- tonydiv 13y agoWe are not making up these numbers. Even for an IO-bound workload, we are seeing 350x speedup. The very fitting application is MapReduce, which is embarrassingly parallel.
- x0x0 13y ago350x on what possible workload? how do you speed up i/o bound workloads -- are you offloading compression/decompression to the gpu or something?
- nemonemo 13y agoThis is a very bold statement. IO-bound means that the bottleneck is at IO. GPU has little benefit in IO. Also, mapreduce is not embarrassingly parallel. Out of mapreduce, map is the only embarrassingly parallel phase. It has input parsing, shuffle, reduction phases that are not embarrassingly parallel.
- Hydraulix989 13y agoI should clarify -- initially I/O bound, but with pipelining, we are able to "impedance match" the flow of data coming off the backing store with the GPU's computational progress. I suppose the better word is "I/O heavy," rather than "I/O bound." We're certainly not the first to describe mapreduce as "embarrassingly parallel," but I can defend myself nonetheless: A shuffle executes in parallel where each unit re-indexes into another unit. Reduction is parallelized over each key, and even finer granularity can be achieved by considering a tree-based approach: http://developer.download.nvidia.com/compute/cuda/1.1-Beta/x86_website/projects/reduction/doc/reduction.pdf http://developer.download.nvidia.com/compute/cuda/1.1-Beta/x... In fact, I used this tree reduction approach to merge non-disjoint clusters of friend groups in my social network clustering algorithm. Parallelizing the input parsing is problem-specific, but even that is possible in many cases.
- 6ren 13y ago(not parent) Yes, since highend GPUs have 1,000's of cores, it makes sense (though I wouldn't expect one GPU core to be faster than one CPU core.) Please, could you quote the CPU and GPU for which you got that 350x speedup? Also, for the 1000x speedup. (also, it would be interesting to know the CPU/GPU that gave the original 1 hour to 0.2 sec (18,000x speedup) - I'm guessing other factors like network latency, low-end CPU + high-end CPU, optimized code etc were part of it.)
- steven2012 13y agoThey mention above they went from Python to presumably C or something else closer to the metal. That probably explains a huge part of the speedup.