3 ms·
That doesn't look so excessive to me. We get hundred or thousand times more efficiency and performance regularly using custom electronics for things like 3d or
by bumbada 5y ago
That doesn't look so excessive to me. We get hundred or thousand times more efficiency and performance regularly using custom electronics for things like 3d or audio recognition.
But programming fixed electronics in parallel is also way harder than flexible CPUs.
"Contemporary core counts coupled with very wide simd makes CPUs functionally similar to ASIC/fpga in many cases."
I don't think so. For things that have a way to be solved in parallel, you can get at least a 100x advantage easily.
There are lots of problems that you could solve in the CPU(serially) that you just can't solve in parallel(because they have inter dependencies).
Today CPUs delegate the video load to video coprocessors of one type or another.
- bumbada 5y agoBTW: Multiple CPUs cores are not parallel programming in the sense fpgas or ASICS (or even GPUs) are. Multiple cores work like multiple machines, but parallel units work choreographically in sync at lower speeds(with quadratic energy consumption). They could share everything and have only the needed electronics that do the job.
- CyberRabbi 5y agoWell transistors are cheap and synchronization is not a bottleneck for embarrassingly parallel video encoding jobs like these. Contemporary CPUs already downclock when they can to save power and conserve heat.
- CyberRabbi 5y ago>> Contemporary core counts coupled with very wide simd makes CPUs functionally similar to ASIC/fpga in many cases. > I don't think so. For things that have a way to be solved in parallel, you can get at least a 100x advantage easily. That’s kind of my point. CPUs are incredibly parallel now in their interface. Let’s say you have 32 cores and use 256 bit simd for 4 64-bit ops. That would give you ~128x improvement compared to doing all those ops serially. It’s just a matter of writing your program to exploit the available parallelism. There’s also implicit ILP going on as well but I think explicitly using simd usually keeps execution ports filled.
- WJW 5y agoTBH 32 or even 64 cores does not sound all that impressive compared to the thousands of cores available on modern GPUs and presumably even more that could be squeezed into a dedicated ASIC. In any case, wouldn't you run out of memory bandwidth long before you can fill all those cores? It doesn't really matter how many cores you have in that case.
- CyberRabbi 5y agoThose thousands of cores are all much more simple and do not have simd and have a huge penalty for branching. There are problems for which GPUs and CPUs are roughly equally well suited. GPUs have their cons.