4 ms·
Interesting, I've recently spent an unhealthy amount of time researching archival formats to build the same setup of using SQLite with ZStd. My use case is ext
by SyrupThinker 3y ago
Interesting, I've recently spent an unhealthy amount of time researching archival formats to build the same setup of using SQLite with ZStd.
My use case is extremely redundant data (specific website dumps + logs) that I want decently quick random access into,
and I was unhappy with either the access speed, quality/usability or even existence of libraries for several formats.
Glancing over the code this seems to use the following setup:
- Aggregate files
- Chunk into blocks
- Compress blocks of fixed size
- Store file to chunk and chunk to block associations
What I did not see is a deduplication step for the chunks, or an attempt to group files (and by extend, blocks) by similarity in an attempt improve compression.
But I might have just missed that due to lack of familiarity with Pascal.
For anyone interested in this strategy, take a look at ZPAQ [1] by Matt Mahoney, you might know him from the Hutter Prize competition [2] / Large Text Compression Benchmark. It takes 14th place with tuned parameters.
There's also a maintained fork called zpaqfranz, but I ran into some issues like incorrect disk size estimates with it.
For me the code was also sometimes hard to read due to being a mix of English and Italian. So your mileage may vary.
[1]: http://mattmahoney.net/dc/zpaq.html http://mattmahoney.net/dc/zpaq.html
[2]: http://prize.hutter1.net http://prize.hutter1.net
[3]: https://github.com/fcorbelli/zpaqfranz https://github.com/fcorbelli/zpaqfranz
- juitpykyk 3y agoFor your use case you might want to look at RocksDB. It supports zstandard compression, random access, and it's very robust.
- OttoCoddo 3y agoThank you for the detail check. I should thank the syrup too :) I'm happy to see a fellow enthusiast. Your deduction is on point. And also, Pack is smart; it skips non-compressible files like MP3 [1], so you do not need to choose the "Store" option to have a faster option, and it speedup decompression too. Pack is the first to achieve this, being faster than Store options. Yes, it was a surprise to me too. ZPAQ is great, and I study the Hutter Prize competition. Pack is on another chart, which is why I proposed CompressedSpeed [2]. The speed of getting to compression needs to be accounted for. You can store anything on an atom if you try hard enough, but hard work takes time. Deduplication step may get added, but in Hard Press [3]. I am curious to see the results of Pack on your data. You can find me here or o at pack.ac. [1] It is based on content rather than extension; any data that is determined not to be worthy of compression, will be stored as is. And as a file can get chucked, some parts can get compressed and some cannot. Imagine that part of the subtitle in a MKV file can get compressed, and the Video part gets skipped. Although these features will get more updates over time, if they don't cost time,. Pack focus is being seamless and not the most compressed; there are already great works in the field, such as the noted ZPAQ. [2] CompressedSpeed = (InputSize / OutputSize) * (InputSize / Speed). Materialized compression speed. [3] You can choose --press=hard to ask for better compression. Even with Hard Press, Pack does not try to eat your hardware just to get a little more; it goes the optimized way I described.
- pdimitar 3y agoTry this one? https://github.com/mhx/dwarfs https://github.com/mhx/dwarfs It has a ton of comparison with existing tools in the README -- zpaqfranz included -- and it seems to be the best there is.