3 ms·
I'm a little bit confused too - all the operations they describe are usually quite fast on a GPU. I wonder if it's not so much about being better at image proc
by TD-Linux 9y ago
I'm a little bit confused too - all the operations they describe are usually quite fast on a GPU.
I wonder if it's not so much about being better at image processing, but having control and direct access to the hardware. For the Adreno, you have to go through Qualcomm's drivers and are subject to their limitations (there is freedreno which works great, but I don't think Qualcomm allows it on handsets). You're also stuck with the sizes of GPU that are available on Snapdragon SoCs - if you want a larger one, too bad.
- makomk 9y agoMost SoCs have some kind of dedicated image processor these days as well as a GPU, even ARM's got in on the game with their own hardware designs. Unfortunately they tend to be pretty proprietary so not much good if you want to do your own image processing. As far as I can tell, the image processors traditionally have a more DSP-like architecture (small local buffers, instruction set optimised for efficient data transfer and processing, etc), but since it's proprietary it's hard to tell. Supposedly they're meant to be more power efficient than the alternatives. (Which isn't surprising; GPUs are really designed for 3D rendering and they're definitely overkill for simpler tasks. The general assumption is that you're fetching small locally-contiguous groups of pixels from main RAM, doing texture mapping and computations on them, then conditionally blitting the result to other locally-contiguous areas of RAM. Most of the infrastructure used for this is going to waste if you're just using it to do 2D image processing.)
- TD-Linux 9y agoIndeed, several good examples of these DSPs are the Qualcomm Hexagon and the VideoCore IV, the latter of which has a 64x64 register bank, and is largely reverse engineered so you can actually figure out how it works [1]. They are really good for highly serial per-block operations, such as video encoding and decoding. They would also work OK for stuff like large convolutions, which is what I kind of imagine the IPUs are for, but not really significantly better than a GPU (the low serial latency is wasted). I did find the original source to the Ars article [2] and it says that each IPU core has 512 ALUs. This seems more like a extremely wide GPU than the VPU. It also seems to be programmable in Halide [3], which presents a more "SIMT"-like interface, just like a GPU shader. It probably lacks the super fast bilinear texture mapping units that a GPU has, but otherwise seems very similar. It'd be interesting to know if the texture/pixel cache is handled automatically, or manually with a large register file like the VPU. The article also features this quote, which also seems to back Google being unsatisfied with only high level access to the GPU: >A key ingredient to the IPU’s efficiency is the tight coupling of hardware and software—our software controls many more details of the hardware than in a typical processor. It would be really awesome if they opened this chip up to third party developers. Unfortunately the press release only mentions availability in the Camera API so far, relegating it to the same opaque blob status as all the other dedicated image processors :( [1] https://github.com/hermanhermitage/videocoreiv/wiki/VideoCore-IV-Programmers-Manual https://github.com/hermanhermitage/videocoreiv/wiki/VideoCor... [2] https://blog.google/products/pixel/pixel-visual-core-image-processing-and-machine-learning-pixel-2/ https://blog.google/products/pixel/pixel-visual-core-image-p... [3] http://halide-lang.org/ http://halide-lang.org/