3 ms·
Rather than doing self-supervised learning on the actual video frames, why not do it on the byte sequence that represents the video file?
by optimalsolver 2y ago
Rather than doing self-supervised learning on the actual video frames, why not do it on the byte sequence that represents the video file?
- mkaic 2y agoYou might find this paper interesting: [JPEG-LM: LLMs as Image Generators with Canonical Codec Representations](https://arxiv.org/abs/2408.08459 https://arxiv.org/abs/2408.08459)
- optimalsolver 2y agoThanks. This is exactly the kind of thing I was looking for.