5 ms·
I wonder if a GPU might be more suitable for his task. Even the 8800GTX is known to do single precision FFTs at more than 55 Gflops which is an order of magnitu
by vizard 18y ago
I wonder if a GPU might be more suitable for his task. Even the 8800GTX is known to do single precision FFTs at more than 55 Gflops which is an order of magnitude more than even contemporary CPUs let alone P4.
- DarkShikari 18y agoOne danger of the term "gigaflop" is how it is measured, and what you use to measure it. Also note that CPUs get a lot better when you start using SIMD code instead of scalar. One classic example of the danger of the word "gigaflop" is that of the exhaustive motion search. If we define a single mathematical operation as a "flop" (technically an iop, since this is integer math), using Sequential Elimination, an optimized exhaustive search algorithm, an 8-core Core 2 system can crank out over 2.7 teraflop-equivalents of processing.
- vizard 18y agoSorry for replying a bit late. For FFTs, flops are measured in a standardized way. If you are doing an FFT of length N, the number of flops is counted as 5 N log N no matter how the actual FFT is computed. So, in the case of FFT, really you just specify a length N and measure the time. For CPUs, the numbers using FFTW, one of the fastest FFT libraries that does take advantage of SIMD, the numbers usually do not exceed 5-6 gflops particularly for larger lengths. OTOH, the above 55 gflops figure is also somewhat misleading since it does not include transfer time of data b/w RAM and GPU. Actual throughput is somewhere around 20gflops. On one particular project using FFT, I got around 15 gflops using GPU including transfer time while testing several FFT libraries, I never got above 3 gflops on a 2.4ghz quad-core using all four cores. The lenghts were big enough not to fit into cache thus reducing CPU performance considerably.