3 ms·
I am not convinced it is eli5. I am too lazy to write a blog post w/ illustration, but for audio signals, which I am more familiar with, the intuition behind DC
by cdavid 6y ago
I am not convinced it is eli5. I am too lazy to write a blog post w/ illustration, but for audio signals, which I am more familiar with, the intuition behind DCT (and MDCT, used e.g. for mp3) is straightforward.
Assuming you understand that a Fourier transform is an operation to go from from time domain to the frequency domain, the problem solved by DCT, DST, etc. is related to the fact that digital signal processing are finite, and without any care, you introduce 'irregularities' if you use a 'normal' Fourier transform.
So the main idea of DCT/DST/etc. is to implicitly 'copy' and/or mirror the signal, to reduce the artefacts/irregularities introduced by Fourier Transform. Reducing irregularities intuitively leads to more regular signals, and the more regular your signal, the quicker the high frequencies decrease, which is the compression effect of DCT.
More mathematically, but still very informally: DCT/DST is about boundary conditions. Using DFT (the 'normal' Fourier transform for digital signals) will imply discontinuities at the boundaries. For continuous time signals, an intuitive way to define regularities is to measure the decay of successive derivative of a function f(t), by looking at convergence convergence of t^n f(t) as t -> inf for n. That implies that regular functions have bounded Fourier transforms, and the more regular, the faster the Fourier transform decays.
The DST/DCT, by mirroring/copying the signal, reduce irregularities, and hence their coefficients decrease faster.
- ponderingfish 6y agoAmazing explanation!!
- jerf 6y agoELI5 was a mistake. ELI12 would have made much more sense. ELI5 is so obviously absurd that people seem to subconsciously just discard the criterion and put all sorts of things in that no conceivable 5-year-olds would understand, not even the genius ones we occasionally read about. ELI5 is extremely restrictive. ELI12 would be much more sensible. An ELI12 explanation would be possible, but it wouldn't be able to casually assume "frequency domain" or related concepts are things that could be assumed; it would first have to show how you can break images down into frequency components.
- cdavid 6y agoYes. I think you can bypass the idea of "frequency domain", and focus on regular / fast changing. I think a smart 12 year old would understand the intuition that fast changing and high frequencies are somehow related, I mean physics gives plenty of concrete example. And then you explain that for images, contours are fast changing, and you have the justification for Fourier coefficients truncation ~ compression. Explaining how DCT helps instead of DFT is the harder part.
- cdavid 6y agoI should add that the OT does not really explain DCT as much as the inherent compressing effect of Fourier transform by removing high frequencies (== what changes quickly, intuitively). That's why the compressed images get blurrier, and also partially why you get more artifacts in JPEG around contours. The problem gets amplified because in images (JPEG) or audio (mp3, etc.), the signal is divided by blocks (windows in audio) first, and those blocks introduce more discontinuities.
- raxxorrax 6y agoBut it doesn't really have an "inherent" compressing effect, only if you opt to remove those frequencies. It is just that the space allows you to quickly determine which information is important or which you can remove. In other cases it is reversible (aside from numerical rounding errors). A square wave signal is a sum of infinite frequencies, but in imaging, the maximum frequency is limited by the resolution of the image.
- cdavid 6y agoI was loose w/ words, but I maintain there is inherent compression aspect to Fourier transform, because for a large class of signals, most of the energy is contained in the first few coefficients. See also the similarity between fourier transforms and KL transform for random variables, itself related to PCA. Ofc, it depends on what you mean exactly by compression. At least for audio, it relates well to keeping the main signal quality because in a sense, the earing system does a frequency decomposition (see e.g. https://www.ncbi.nlm.nih.gov/books/NBK10946/figure/A896/?report=objectonly https://www.ncbi.nlm.nih.gov/books/NBK10946/figure/A896/?rep...), and the higher the frequency, the less sensitive to it we tend to be.
- Jasper_ 6y agoRight. The DCT, in theory, is completely lossless. But the reason that the "other space" is helpful for us for compression is because our eyes and ears are more interested in comparing relatively than absolutely. So you don't need the exact high-frequency content, just the relatively correct high-frequency content. As discussed, "high-frequency" just means stuff that's moving very fast. In sound, that's compared to pitch. In a JPEG image, that means a given pixel's brightness. This is what the DCT is exploiting -- that we can adjust the high-frequency terms and it will still look OK to our eyes. https://magcius.github.io/dctx/ https://magcius.github.io/dctx/
- tspiteri 6y agoI would also mention something about that with both the DFT and the DCT, the signal is conceptually periodic in time. So if I want to find the DFT of the 4-sample sequence 1, 2, 3, 4, it is like finding the DFT of the sequence 1, 2, 3, 4, 1, 2, 3, 4, 1, 2, ..., which has the discontinuity you mention between the 4s and 1s. With the DCT, it's like the periodic sequence is 1, 2, 3, 4, 4, 3, 2, 1, 1, 2,..., which has less discontinuity. (There are different DCT variants on the exact offset of the reflection, but that can be ignored in this discussion.)