4 ms·
Why not use the GPU? This is exactly the kind of tasks GPUs were designed for. Gimp can do this in real time using the difference layer mode.
by ajfjrbfbf 6y ago
Why not use the GPU? This is exactly the kind of tasks GPUs were designed for. Gimp can do this in real time using the difference layer mode.
- detaro 6y agoDoes using the GPU pay off for one-off diffing? It's "obvious" that it makes sense for Gimp, but does it for this?
- ben-schaaf 6y agoCompared to a threaded CPU implementation, it only pays off if your image is already sitting in VRAM. This kind of operation should be primarily limited by memory bandwidth, so you're just making it take longer by copying the image to VRAM and copying the diff back.
- formerly_proven 6y agoThe bottleneck is almost certainly decoding the image (PNGs are zlib compressed, which will be going at something like ~300 MB/s plus the filter stuff PNG does), not comparing raw pixels (which you should be able to do at memory bandwidth, so something in excess of 10 GP/s).
- eps 6y agoCompressed, but not as a single continuous block. The bottleneck will probably be in doing things serially, both with decompression and the disk IO.
- viraptor 6y agoEven before you get into GPU, it looks like you could do the normalisation here https://github.com/n7olkachev/imgdiff/blob/master/pkg/yiq/delta.go#L13 https://github.com/n7olkachev/imgdiff/blob/master/pkg/yiq/de... using _mm_mul_ps or _mm256_mul_ps depending on availability instead of standard math. (CPU based vector processing) There's also bound to be some vector compare trick as well which falls back to pixel-by-pixel only for differences. I think there are quite a few simple tricks you could apply here to get a bit more performance... Starting with sacrificing some memory and decoding a chunk of image at a time rather than calling .At(x,y) every time. And maybe improving local cache by dividing the image in `cpu` parts rather than having the work interleaved.
- kaetemi 6y agoConsidering the source material is a file, that needs to be loaded up by the CPU first, and the operation itself is dead simple, the real constraint is RAM bandwidth. I would guess that the time of getting the data out to the GPU and back, is about the same as just calculating the difference on the CPU. So, the quickest solution is probably to use the CPUs vector operations, and read ahead into the registers so you can exhaust the memory bandwidth fully.
- amelius 6y agoI recently tried this with a 19 megabyte .tiff file in PyTorch. Uploading the tensor took (i think) around 1 or 2 seconds, wall-clock time. So I quickly dismissed the GPU solution to my problem. But I'm still wondering... is this normal?
- kaetemi 6y agoYou've got your GPU program that needs to be compiled to actual GPU code at startup as well. That takes a bit. The actual file transfer is comparable to the GPU doing a memcpy. (You could also use pinned host memory if the GPU program is just doing a one-off read.) GPU processing for simple operations really shines if your data is already on the GPU and your code is hot. Otherwise, it's more useful if the process is significantly complex enough to make the overhead of a cold startup insignificant.
- Kuinox 6y agoGPU can now directly read the file on the SSD.
- plasticchris 6y agoCitation needed? I know things have been announced but what's available now?
- harias 6y agoGPUDirect has been available for quite sometime now: https://developer.nvidia.com/gpudirect https://developer.nvidia.com/gpudirect