3 ms·
I should add that the OT does not really explain DCT as much as the inherent compressing effect of Fourier transform by removing high frequencies (== what chang
by cdavid 6y ago
I should add that the OT does not really explain DCT as much as the inherent compressing effect of Fourier transform by removing high frequencies (== what changes quickly, intuitively). That's why the compressed images get blurrier, and also partially why you get more artifacts in JPEG around contours.
The problem gets amplified because in images (JPEG) or audio (mp3, etc.), the signal is divided by blocks (windows in audio) first, and those blocks introduce more discontinuities.
- raxxorrax 6y agoBut it doesn't really have an "inherent" compressing effect, only if you opt to remove those frequencies. It is just that the space allows you to quickly determine which information is important or which you can remove. In other cases it is reversible (aside from numerical rounding errors). A square wave signal is a sum of infinite frequencies, but in imaging, the maximum frequency is limited by the resolution of the image.
- cdavid 6y agoI was loose w/ words, but I maintain there is inherent compression aspect to Fourier transform, because for a large class of signals, most of the energy is contained in the first few coefficients. See also the similarity between fourier transforms and KL transform for random variables, itself related to PCA. Ofc, it depends on what you mean exactly by compression. At least for audio, it relates well to keeping the main signal quality because in a sense, the earing system does a frequency decomposition (see e.g. https://www.ncbi.nlm.nih.gov/books/NBK10946/figure/A896/?report=objectonly https://www.ncbi.nlm.nih.gov/books/NBK10946/figure/A896/?rep...), and the higher the frequency, the less sensitive to it we tend to be.
- Jasper_ 6y agoRight. The DCT, in theory, is completely lossless. But the reason that the "other space" is helpful for us for compression is because our eyes and ears are more interested in comparing relatively than absolutely. So you don't need the exact high-frequency content, just the relatively correct high-frequency content. As discussed, "high-frequency" just means stuff that's moving very fast. In sound, that's compared to pitch. In a JPEG image, that means a given pixel's brightness. This is what the DCT is exploiting -- that we can adjust the high-frequency terms and it will still look OK to our eyes. https://magcius.github.io/dctx/ https://magcius.github.io/dctx/