3 ms·
The first draft of the article actually had that reason, but there is also a strong correlation between the size of the dict (these dicts are almost 1Mb, while
by SaveTheRbtz 10y ago
The first draft of the article actually had that reason, but there is also a strong correlation between the size of the dict (these dicts are almost 1Mb, while other languages are closer to 500kb) and compression ratio improvements. Therefore I've played it safe and attributed it to the window size.
Though for languages like Korean and Chinese (whose size is more inline with latin languages) we see 27.5% improvement, which is most likely due to context modeling.
Therefore I assume ratio improvement is split ~50/50 between these two. It was easy to verify that by compressing data with `brotli --window 15` and comparing ratios there, but I was lazy there. I'm sorry.
PS. I've also skipped NFC/NFD part of the post which is very interesting for Korean, where NFC normalized text occupies 30% less space. It also gives additional ratio 5% for brotli and 15% for gzip.
- JyrkiAlakuijala 10y agoThere is no Thai or Korean in the dict. The total size of the dict (including all languages) is 120 kB.
- SaveTheRbtz 10y agoBy "dict" I meant the data we are compressing: these are basically dictionaries for "English to X" translation. What was saying is that there is a strong correlation between size of the data I was compressing and compression ratio improvements over gzip.