4 ms·
I would worry that the fixed, non-overlapping block nature of a JPEG would reduce translation invariance - shift an image by 4 pixels and the DCT coefficients m
by johntb86 3y ago
I would worry that the fixed, non-overlapping block nature of a JPEG would reduce translation invariance - shift an image by 4 pixels and the DCT coefficients may look very different. People have been doing a lot of work to try to reduce the dependence of the image on the actual pixel coordinates - see for example https://research.nvidia.com/publication/2021-12_alias-free-generative-adversarial-networks https://research.nvidia.com/publication/2021-12_alias-free-g...
- DougBTX 3y agoOn the other hand, ViT uses non-overlapping patches anyway, so the impact may be minor. Example code: https://nn.labml.ai/transformers/vit/index.html https://nn.labml.ai/transformers/vit/index.html
- andai 3y agoDoes JPEG-2000 fix that? From what I gathered, it doesn't use blocks.
- mochomocha 3y agoJPEG-2000 uses wavelets as a decomposition basis as opposed to DCT which in theory makes it possible to treat the whole image as a single block while ensuring high compression. In practice though tiles are used, I would guess to improve on memory and compute parallelism.