4 ms·
That comparison can not be directly made as the 'O' in "FLOPS" is different.
by quickben 7y ago
That comparison can not be directly made as the 'O' in "FLOPS" is different.
- Const-me 7y agoThey're math operations on small vectors of 32-bit floating point numbers, they produce same result given same input data. The only difference is that CPU computes stuff on 4-wide vectors of these numbers (modern CPUs up to 16), that GPU on 32-wide vectors (other modern GPUs up to 64).
- Dylan16807 7y agoA lot of the FLOPS on newer CPUs are like GPUs, yes. But the comparison is against an old CPU, without big vector units. Those non-vector calculations are very different, and far more flexible. A brand new CPU core can do 3-4x as many separate operations per cycle, and is clocked 6-10x as fast. Having more cores helps but it's still far behind. Also for a fair price/performance ratio you probably want to compare to the 450MHz model at $230, so only a $350 CPU today ($280 equivalent by august). https://money.cnn.com/1999/08/23/technology/intel/ https://money.cnn.com/1999/08/23/technology/intel/
- Const-me 7y ago> and far more flexible They’re both very flexible, just different. For example, GPUs are better at vectorized condition code. CPUs were mostly fixed with AVX512 but these instructions are too new, only available on some servers. Sure, there’re algorithms which don’t work on GPUs. A stream cipher would be very slow because requires single-thread performance, also GPUs don’t have AES hardware. A compiler is borderline impossible because inherently scalar and requires dynamic memory. Also GPUs don’t do double-precision math particularly fast. Still, I think many users who need high-performance computing can utilize GPUs. They’re trickier to program, but this might be fixable with tools/languages, we have been programming classic computers for ~70 years, GPGPUs for just 12.
- Dylan16807 7y ago> For example, GPUs are better at vectorized condition code. That's throughput, not flexibility. I would define flexibility in terms of how easily the instruction stream can vary per math operation. Full flexibility requires a lot more transistors per FLOP, which is why you can't use wildly different architectures to assess Moore's law, which is about transistor count. And comparing transistors on Pentium 3 (including cache) to an RTX 2060 (including cache) it seems to be 34 million vs. 10800 million. That's two and a half orders of magnitude.