3 ms·
Right, and this why I believe that FPGAs aren't competitive with GPUs for computationally-bound tasks. FPGAs are weak here because (a) routing costs power, (b)
by ianhowson 7y ago
Right, and this why I believe that FPGAs aren't competitive with GPUs for computationally-bound tasks.
FPGAs are weak here because (a) routing costs power, (b) Most ALUs need to be implemented from LUTs, (c) high-order data flow needs to consider timing of all dependencies, manually, and getting this wrong is very wasteful of resources.
GPUs have an apparent weakness where their ALUs are overly general for most tasks. If you're doing primarily INT8 tasks then all of that floating point hardware goes to waste, and it appears that an FPGA has an advantage (INT8 is cheap and predictable there.)
However, ALU volume isn't the constraint on GPUs -- it's memory bandwidth. Every prosumer GPU make in the last decade has had an excess of ALUs and a shortage of memory ops per second.
This is where we're seeing specialized TPUs, deep learning accelerators and image processing accelerators that provide the right ops paired with the right types of memory in the right places. They're seeing large performance and power improvements over GPUs, at the cost of being further specialized.
FPGAs have an advantage for low-latency applications and where predictable timing (nanoseconds) is required, but I don't see them displacing GPUs for compute-bound tasks.
- shaklee3 7y agoThe Volta and higher architecture share ALUs with the floating point units, so nothing is really going to waste. Same with the tensor cores.
- imtringued 7y agoThe weakness of GPUs is that the minimum amount of work is incredibly high for acceleration to make sense. If you have a batch with less than 10000 units of work then you shouldn't even think of writing a GPU kernel. GPUs will never used for stream/network processing workloads because no one wants to wait for a buffer to fill up. A FPGA could in theory offer the same throughput but with a much lower latency.