3 ms·
> ignore the cost of initial weights Well, then Wikipedia itself is a very good compression that only needs the title to perfectly predict the full article.
by JohnKemeny 2mo ago
> ignore the cost of initial weights
Well, then Wikipedia itself is a very good compression that only needs the title to perfectly predict the full article.
- adamgordonbell 2mo agoExcept the LLM generalizes its encoding to all english text where as the copy of wikipedia can only 'compress' wikipedia.
- JohnKemeny 2mo agoI probably didn't understand what you meant by > Hutter Prize being where you are paid if you can compress wikipedia small enough. LLMs do very well at that, if, big if, you ignore the cost of initial weights. then. If all you care about is compressing Wikipedia, but ignore the size of the actual data, what is it that you are actually trying to do?