4 ms·
> It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine le
by zerd 1y ago
> It has long been established that predictive models can be transformed into lossless
compressors and vice versa. Incidentally, in recent years, the machine learning
community has focused on training increasingly large and powerful self-supervised
(language) models. Since these large language models exhibit impressive predictive
capabilities, they are well-positioned to be strong compressors. In this work, we
advocate for viewing the prediction problem through the lens of compression and
evaluate the compression capabilities of large (foundation) models. We show
that large language models are powerful general-purpose predictors and that the
compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively.
https://arxiv.org/pdf/2309.10668 https://arxiv.org/pdf/2309.10668
Transformers are also used in the top algorithm right now on the Large Text Compression Benchmark. https://bellard.org/nncp/nncp.pdf https://bellard.org/nncp/nncp.pdf