4 ms·
The point you make is really valuable. A few days ago, I found myself explaining to someone why custom hardware and GPUs could so easily outperform processors,
by PieSquared 14y ago
The point you make is really valuable. A few days ago, I found myself explaining to someone why custom hardware and GPUs could so easily outperform processors, and I realized that most programmers have no concept of how much overhead the general nature of a processor entails. (Although, I don't think most programmers really need to know this.)
For instance, let's take the problem of multiplying ten numbers. In a normal processor, you have a loop of instructions, each instruction has to go through a "fetch" state (to load it from memory), a "decode" stage, to figure out what the instruction is, an "issue" stage, to figure out which processor pipeline can best execute this instruction, an "execute" stage, to finally execute the instruction, and maybe a "commit" stage to write the outputs back to memory. (The exact number of stages and amount of parallelism depends on the microarchitecture and pipeline depth, of course). What if we wanted to just build a chip that did this? We could put ten multipliers on the chip, and then do the exact same operations in just a few clock cycles, since we would have no instruction fetch or decode, no commit, no loops, and so on. This is a contrived example, but my point is that general-purpose processors are incredibly slow compared to dedicated hardware, precisely because the extra transistors necessary to make processors general purpose also take a large portion of the computing time.
I find the idea of FPGAs reconfigured per-application to be really interesting. Celoxica (http://www.celoxica.com http://www.celoxica.com) seems to do some sort of FPGA-based software acceleration for trading software, for instance. I wonder if it's possible to do something like this for a more general market...