5 ms·
If you’re compressing 100 or 100k such datasets, presuming that it is not custom tuned for this corpus, then wouldn’t you still save much more than you spend?
by binary132 2y ago
If you’re compressing 100 or 100k such datasets, presuming that it is not custom tuned for this corpus, then wouldn’t you still save much more than you spend?
- remram 2y agoI'm not saying the result is completely useless, I am comparing it to the age-old technique of using a dictionary. Does this new LLM-powered technique improve upon the old dictionary technique? Dictionaries also don't require a GPU or this amount of RAM. Where I assume LLMs would shine is lossy compression.
- binary132 2y agoAh ok, I think we made different assumptions about whether the model was specific to the particular dataset so each one would need a new model — a dictionary is specific to the particular dataset being compressed, right? I was thinking the LLM would be a general-purpose text compression model.
- remram 2y agoNot particularly. You could make a dictionary from "the English web", with common character sequences found on those sites you use as input.
- ksec 2y agoI have the same question, what is the different between LLM and Dictionary in the context of compression. Can I not "train" a dictionary?
- binary132 2y agoAIUI, a dictionary is built during compression to specify the heuristics of a particular dataset and belongs to that specific dataset only. For example, it could be a ranking of the most frequent 10 symbols in the compressed file. That will be different for every input file.
- mbreese 2y ago> That will be different for every input file That could be different for every input file, but it doesn't have to be. It could also be a fixed dictionary. For example, ZLIB allows for a user-defined dictionary [1]. In this case, I'd consider the LLM to be a fixed dictionary of sorts. A very large, fixed dictionary with probabilistic return values. [1] https://www.rfc-editor.org/rfc/rfc1950#page-9 https://www.rfc-editor.org/rfc/rfc1950#page-9
- binary132 2y agoAh, I see. I’d never thought of the possibility of using a dictionary not created specifically from the given input dataset, heh
- mbreese 2y agoAdmittedly, I don’t think it is common, but I think there was a project a few years ago (Google?) that tried to compress HTML using at least a partially fixed dictionary. Nowadays though, it’s apparently still something that’s being tried. Chrome now supports shared dictionaries for Zstd and Brotli. One idea being, you would likely benefit from having a shared dictionary used to decompress multiple artifacts for a site. But, you many not want everything compressed all together, so this way you get the compression benefit, but can have those artifacts split into different files. https://developer.chrome.com/blog/shared-dictionary-compression https://developer.chrome.com/blog/shared-dictionary-compress...