4 ms·
One important intuition about gzip vs snappy is that gzip isn’t generally parallelizable, while snappy is. If your Hadoop cluster is storing enormous files this
by Godel_unicode 4y ago
One important intuition about gzip vs snappy is that gzip isn’t generally parallelizable, while snappy is. If your Hadoop cluster is storing enormous files this shows up very quickly in read performance.
- mxmlnkn 4y agoDecompression of arbitrary gzip files can be parallelized with pragzip: https://github.com/mxmlnkn/pragzip https://github.com/mxmlnkn/pragzip
- Godel_unicode 4y agoThus the words “isn’t generally” as opposed to “cannot be”.
- glogla 4y agoThat however only impacts CSVs or other text files, where the compression is the envelope. In parquet, it is the internal chunks that are compressed, so it is as parallelizable as ever. Also, generic CSV is not parallelizable anyway, because it allows you to have (quoted) newlines inside fields and then there is really now way to split it on newlines because you never know if you're on beginning or row or in the middle. So parallel reading of CSVs is more of a optimization for specific CSVs that you know don't have newlines inside fields.