4 ms·
Vision transformers are essentially just JPEG but with learned features rather than the Fourier transform.
by programjames 2y ago
Vision transformers are essentially just JPEG but with learned features rather than the Fourier transform.
- larodi 2y agoIndeed. Kernels mashing features. Knowing jpeg helped the understanding of embedding a lot. It’s why I tell friends - talking to GPT is like talking to .ZIP files…
- Zacharias030 2y agoInteresting! Can you elaborate?
- programjames 2y agoThe JPEG algorithm is: 1. Divide up the image into 8x8 patches 2. Take the DCT (a variant of the Fourier transform) of each patch to extract key features 3. Quantize the outputs 4. Use arithmetic encoding to compress The ViT algorithm is: 1. Divide up the image into 16x16 patches 2. Use query/key/value attention matrices to extract key features 3. Minimize cross-entropy loss between predicted and actual next tokens. (This is equivalent to trying to minimize encoding length.) ViT don't have quantization baked into the algorithm, but NNs are being moved towards quantization in general. Another user correctly pointed out that vision transformers are not necessarily autoregressive (i.e. they may use future patches to calculate values for previous patches), while arithmetic encoding usually is (so JPEG is), so the algorithms have a few differences but nothing major. ----- I think it's pretty interesting how closely related generation and compression are. ClosedAI's Sora[^1] model uses a denoising vision transformer for their state-of-the-art video generator, while JPEG has been leading image compression for the past several decades. [^1]: https://openai.com/index/sora/?video=big-sur https://openai.com/index/sora/?video=big-sur
- ImageXav 2y agoI think it's important to point out for people that might be interested in this comment that a few things are wrong. 1. Standard JPEG compression uses the Discrete Cosine Transform, not the Fourier Transform. 2. It is easy to be dismissive of any technology by saying that it is 'just' X with Y, Z, etc on top 3. Vision transformers allow for much longer range context - the magic comes in part from the ability to relate between patches, as well as the learned features, which JPEG does not do.
- programjames 2y agoThe discrete cosine transform is the real part of a Fourier transform.
- idiotsecant 2y agoRockets are essentially just fire that burns real fast.