4 ms·
I found parquet files to be slower than just serializing and compressing the pandas data frames. Haven't looked back since. Of course this was an older version
by cheez 6y ago
I found parquet files to be slower than just serializing and compressing the pandas data frames. Haven't looked back since. Of course this was an older version of pandas so the DF to parquet functionality may be much improved.
- mumblemumble 6y agoA serialization format for Pandas isn't really the core use case of Parquet. Pandas will slurp the whole thing into memory, which doesn't take advantage of the columnar format or pushdown filtering features. It gets more interesting if you use it with a tool that uses them to avoid some large fraction of disk I/O when performing a selective query.
- cheez 6y agoI just didn't trust the tools to do the job properly so for now I just split it up myself.
- mumblemumble 6y agoConsidering the number of segfaults I've witnessed that appear to originate in relatively recent versions of the Parquet library, I don't think your instincts are misplaced. I genuinely worry about data corruption. I've been meaning to take a closer look at ORC. As a Spark user, I just sort of defaulted into Parquet. ORC is very similar, though, and seemingly gives every indication of being the more mature product.
- cheez 6y agoOn balance I find writing your own tools is useful when your use cases are narrow but becoming a wizard in other tools is useful otherwise. I'm lucky that I get to define my own use cases.