3 ms·
That depends how much text you're compressing. For example, LLM pretraining datasets are on the order of tens or hundreds of terabytes. If an LLM-based code fo
by davmre 2mo ago
That depends how much text you're compressing.
For example, LLM pretraining datasets are on the order of tens or hundreds of terabytes. If an LLM-based code for that data is ~twice as efficient as gzip, you could afford to transmit the weights of even a very large LLM and still come out ahead.
In other words: LLMs actually are excellent compressors of their training sets in the formal information-theoretic sense.