4 ms·
The article [1] is a bit old (2010) but still relevant to answer your question. It compares the efficiency of a HD.264 implemented first on a general purpose CP
by yaantc 4y ago
The article [1] is a bit old (2010) but still relevant to answer your question. It compares the efficiency of a HD.264 implemented first on a general purpose CPU, then on a SIMD+VLIW DSP, then on the same DSP with additional optimized instruction, then on this platform with macro memory based accelerators, and lastly a pure ASIC implementation. There are significant efficiency gains at each steps, the big gap is between optimized ops and macro accelerators.
Flexibility do have a price. With efficiency gains due to new nodes getting lower (and pricer!), it's normal to see more dedicated hardware accelerators.
[1] https://courses.cs.washington.edu/courses/cse591n/10au/Papers/Hameed2010_SourcesOfInefficiency.pdf https://courses.cs.washington.edu/courses/cse591n/10au/Paper...