6 ms·
.parquet files are completely underrated, many people still do not know about the format! .parquet preserves data types (unlike CSV) They are 10x smaller than
by mrtimo 2y ago
.parquet files are completely underrated, many people still do not know about the format!
.parquet preserves data types (unlike CSV)
They are 10x smaller than CSV. So 600GB instead of 6TB.
They are 50x faster to read than CSV
They are an "open standard" from Apache Foundation
Of course, you can't peek inside them as easily as you can a CSV. But, the tradeoffs are worth it!
Please promote the use of .parquet files! Make .parquet files available for download everywhere .csv is available!
- sph 2y agoThird consecutive time in 86 days that you mention .parquet files. I am out of my element here, but it's a bit weird
- fifilura 2y agoFWIW I am the same. I tend to recommend BigQuery and AWS/Athena in various posts. Many times paired with Parquet. But it is because it makes a lot of things much simpler, and that a lot of people have not realized that. Tooling is moving fast in this space, it is not 2004 anymore. His arguments are still valid and 86 days is a pretty long time.
- ok_computer 2y agoSometimes when people discover or extensively use something they are eager to share in contexts they think are relevant. There is an issue when those contexts become too broad. 3 times across 3 months is hardly astroturfing for big parquet territory.
- mrtimo 2y agoI've downloaded many csv files that were mal-formatted (extra commas or tabs etc.), or had dates in non-standard formats. Parquet format probably would not have had these issues!
- swyx 2y agono need to be so suspicious when its an open standard not even linked to a startup?
- ddalex 2y agoWhy is .parquet better than protobuf?
- sdenton4 2y agoParquet is columnar storage, which is much faster for querying. And typically for protobuf you deserialize each row, which has a performance cost - you need to deserialize the whole message, and can't get just the field you want. So, of you want to query a giant collection of protobufs, you end up reading and deserializing every record. For parquet, you get much closer to only reading what you need.
- ddalex 2y agoThank you.
- nostrademons 2y agoParquet ~= Dremel, for those who are up on their Google stack. Dremel was pretty revolutionary when it came out in 2006 - you could run ad-hoc analyses in seconds that previously would've taken a couple days of coding & execution time. Parquet is awesome for the same reasons.
- thesz 2y agoParquet is underdesigned. Some parts of it do not scale well. I believe that Parquet files have rather monolithic metadata at the end and it has 4G max size limit. 600 columns (it is realistic, believe me), and we are at slightly less than 7.2 millions row groups. Give each row group 8K rows and we are limited to 60 billion rows total. It is not much. The flatness of the file metadata require external data structures to handle it more or less well. You cannot just mmap it and be good. This external data structure most probably will take as much memory as file metadata, or even more. So, 4G+ of your RAM will be, well, used slightly inefficiently. (block-run-mapped log structured merge tree in one file can be as compact as parquet file and allow for very efficient memory mapped operations without additional data structures) Thus, while parqet is a step, I am not sure it is a step in definitely right direction. Some aspects of it are good, some are not that good.
- datadeft 2y agoNobody is forcing you to use a single Parquet file.
- thesz 2y agoOf course. But nobody tells me that I can hit a hard limit and then I need a second Parquet file and should have some code for that. The situation looks to me as if my "Favorite DB server" supports, say, only 1.9 billions records per table and if I hit that limit I need a second instance of my "Favorite DB server" just for that unfortunate table. And it is not documented anywhere.
- apwell23 2y agosome critiques of parquet by andy pavlo https://www.vldb.org/pvldb/vol17/p148-zeng.pdf https://www.vldb.org/pvldb/vol17/p148-zeng.pdf
- thesz 2y agoThanks, very insightful. "Dictionary Encoding is effective across data types (even for floating-point values) because most real-world data have low NDV ratios. Future formats should continue to apply the technique aggressively, as in Parquet." So this is not critique, but assessment. And Parquet has some interesting design decisions I did not know about. So, let me thank you again. ;)
- riku_iki 2y ago> They are 50x faster to read than CSV I actually benchmarked this and duckdb CSV reader is faster than parquet reader.
- wenc 2y agoI would love to see the benchmarks. That is not my experience, except in the rare case of a linear read (in which CSV is much easier to parse). CSV underperforms in almost every other domain, like joins, aggregations, filters. Parquet lets you do that lazily without reading the entire Parquet dataset into memory.
- riku_iki 2y ago> That is not my experience, except in the rare case of a linear read (in which CSV is much easier to parse). Yes, I think duckdb only reads CSV, then projects necessary data into internal format (which is probably more efficient than parquet, again based on my benchmarks), and does all ops (joins, aggregations) on that format.
- wenc 2y agoYes, it does that, assuming you read in the entire CSV, which works for CSVs that fit in memory. With Parquet you almost never read in the entire dataset and it's fast on all the projections, joins, etc. while living on disk.
- riku_iki 2y ago> which works for CSVs that fit in memory. what? Why CSV is required to fit in memory in this case? I tested CSVs which are far larger than memory, and it works just fine.
- geysersam 2y agoThe entire csv doesn't have to fit in memory, but the entire csv has to pass through memory at some point during the processing. The parquet file has metadata that allows duckdb to only read the parts that are actually used, reducing total amount of data read from disk/network.
- jjgreen 2y agoPlease promote the use of .parquet files! apt-cache search parquet <nada> Maybe later
- seabass-labrax 2y agoParquet is a file format, not a piece of software. 'apt install csv' doesn't make any sense either.
- jjgreen 2y agoThere is no support for parquet in Debian, by contrast apt-cache search csv | wc -l 259
- fhars 2y agoIf you want to shine with snide remarks, you should at least understand the point being made: $ apt-cache search csv | wc -l 225 $ apt-cache search parquet | wc -l 0
- riku_iki 2y agoapt search would return tons of libparquet-java/c/python packages if it was popular.
- nostrademons 2y agoIt's more like "sudo pip install pandas" and then Pandas comes with Parquet support.
- jjgreen 2y agoPandas cannot read parquet files itself, it uses 3rd party "engines" for that purpose and those are not available in Debian
- nostrademons 2y agoAh yes, that's true though a typical Anaconda installation will have them automatically installed. "sudo pip install pyarrow" or "sudo pip install fastparquet" then.
- swyx 2y ago> They are 10x smaller than CSV. So 600GB instead of 6TB. how? lossless compression? under what scenario? vague headlines like this just beg more questions
- riku_iki 2y agoLikely this assumes that parquet has internal compression applied, and CSV is uncompressed.
- alentred 2y agoAgreed. The abstractions on top of parquet are quite immature yet, though, and lots of software assumes that if you use Parquet - you also use Hive, Spark and stuff. Take Apache Iceberg for example. It is essentially a specification to how to store parquet files for efficient use and exploration of data, but the only implementation... depends on Apache Spark!
- 62951413 2y agoYou kind of can peek into parquet files with a tiny command line utility: https://github.com/manojkarthick/pqrs https://github.com/manojkarthick/pqrs