5 ms·
(OP here) I actually quite agree with you on all of these points. There are tools, and programming techniques, for dealing with everything in my rant. And I als
by memset 2y ago
(OP here) I actually quite agree with you on all of these points. There are tools, and programming techniques, for dealing with everything in my rant. And I also agree that many times people are "holding it wrong" - be it data that has not been partitioned, or people trying to do a "SELECT *" against a terabyte of parquet files.
My central point is that the "bucket of parquet" can be wrangled with by an expert, who knows the tools, their limitations, and programming techniques, in order to get it into a usable state.
But my post is a frustration with the sheer amount of tooling and knowledge required to get started - for example, the case where non-partitioned data is foisted upon non-experts.
- organsnyder 2y ago> But my post is a frustration with the sheer amount of tooling and knowledge required to get started - for example, the case where non-partitioned data is foisted upon non-experts. How is that different from throwing a massive relational database at someone who doesn't know how to manage indexes and other optimizations?
- bradford 2y agoNot OP, but I'd guess there's greater industry awareness of relational DBs than there are of parquet files. I've been on the receiving end of a Parquet file that I didn't know how to crack open the ambiguity on how to proceed was frustrating.
- wenc 2y agoThis is true. There are two tools you need to know for this: duckdb and visidata. With these tools, Parquet is almost as easy as CSVs (but a few orders of magnitude more powerful and faster) Parquet is also usable in polars and pandas, and Apache Spark too but that’s getting into complicated territory. DuckDB it’s literally just Select * from ‘s3://bucket/*.parquet’
- add-sub-mul-div 2y agoRelational database technology is more highly proven, stable, documented, and consistent than the constellation of big data solutions. Learning about indexes etc. 20 years ago would still help you today. Learning about this year's big data stack may not even help you next year.
- wenc 2y agoIf you’re dealing with non trivial data sizes and if you have an analytics use case and you’re using CSV, you’re already doing it wrong. You bring up partitioning — that can help performance but with Parquet, you can get performance they is far better than CSV even without partitioning because it only has to read the headers. And no, no one really needs to know about row groups (sure you can eke out more performance but few people do). All this is to say all your points are non starters. There’s no need to know all this and even the least optimized parquet dataset is better than CSV in every way. The only use case CSV might be good for is ETL into an actual database like Postgres which is what you’re doing in your article but that’s actually only a very small part of the analytics pipeline. Parquet is actually not that complicated but I always meet people who only know CSV and they feel anything more complicated is beyond them. Don’t be that guy. Shoehorning everything into CSV creates costs downstream like high storage costs (CSVs are much bigger), and no type safety means you have to validate types with a bunch of non standard rules creating a lot of technical debt. I’ve always managed to work with parquet with either DuckDB or Visidata (and occasionally Pandas or Polars). Visidata lets you peek into parquet like Vim lets you look into CSVs. I don’t miss CSVs at all.
- int_19h 2y agoThe article doesn't claim that CSV is superior, though. It only mentions CSV once, as a workaround for the problem that relational databases mostly don't support importing from Parquet directly. It does rather feel like at least that part of the rant would be better addressed by adding such support, though. If Parquet is already a de facto standard in the industry, I'm sure there already are Postgres extensions that add such support; how hard would it be to convince the maintainers to take the best-written one and use it to add support directly to core Postgres, just as it already has for CSV and JSON?