3 ms·
CPUs and GPUs are fundamentally aiming at different vertices of the performance polygon. CPUs aim to minimize latency (how many cycles have to pass before you
by owlbite 2y ago
CPUs and GPUs are fundamentally aiming at different vertices of the performance polygon.
CPUs aim to minimize latency (how many cycles have to pass before you can use a result), and do so by way of high clock frequencies, caches and fancy micro architectural tricks. This is what you want in most general computation cases where you don't have other work to do whilst you wait.
GPUs instead just context switch to a different thread whilst waiting on a result. They hide their latency by making parallelism as cheap as possible. You can have many more cores running at a lower clock frequency and be more efficient as a result. But this only works if you have enough parallelism to keep everything busy whilst waiting for things to finish on other threads. As it happens that's pretty common in large matrix computations done in machine learning, so they're pretty popular there.
Will they converge? I don't think so - they're fundamentally different design points. But it may well be that they get integrated at a much closer level than current designs, pushing the heterogeneous/dark silicon/accelerator direction to an extreme.