3 ms·
In most real-world cases, it'll look like this: gzip+Homogenius < gzip < Homogenius < original JSON gzip+Homogenius will typically be better than just gzip
by walrus 12y ago
In most real-world cases, it'll look like this:
gzip+Homogenius < gzip < Homogenius < original JSON
gzip+Homogenius will typically be better than just gzip because Homogenius 'knows' about the structure of JSON. gzip will typically be better than Homogenius (without gzip) because it can compress the payload.
Here's the size (in bytes) of the sample files given with the project:
69934 packed-homogenius.json.gz
79550 verbose.json.gz
213742 packed-homogenius.json
265177 verbose.json
- beefsack 12y agoYou had the same idea as me, I think I posted my edit almost exactly when you posted this! Interesting to see a similar correlation in your results in the ratio between the two non-gzipped and gzipped files.
- _delirium 12y agoI get the same ordering with gzip, but it seems sensitive to compression algorithm. With bzip2 (default settings, v1.0.6), homogenius seems to actually worsen the compression ratio, at least on this one example: 52093 verbose.json.bz2 56298 packed-homogenius.json.bz2 68521 packed-homogenius.json.gz 79222 verbose.json.gz 213742 packed-homogenius.json 265177 verbose.json
- Scaevolus 12y agoBWT-based compressors like bzip2 do best when their inputs have highly repetitive structure. In a JSON file with many repeated keys, the information that `"k` is usually followed by `ey_1": "` is compressed very effectively. For similar reasons, bzip2 tends to outperform gzip on executables-- it can more effectively model the conditional probability of opcode sequences. See http://mattmahoney.net/dc/dce.html#Section_55 http://mattmahoney.net/dc/dce.html#Section_55 for more discussion