3 ms·
The article isn't particularly about Parquet or S3. They are both great technologies and I use them all the time. I mentioned CSV because it is the only format
by memset 2y ago
The article isn't particularly about Parquet or S3. They are both great technologies and I use them all the time.
I mentioned CSV because it is the only format supported by Postgres for ingest (besides their binary protocol) but that is only an instance of the general problem of interoperability.
I'm not sure I agree that "reading a ton of data is not meant to be easy." I'd say it is not easy today, and there is a constellation of tools and programming techniques that can let you perform the task, but it is often a distraction from the end result you want to achieve.
The essence of my rant is that there are many steps distracting steps between "a data dump" and "the thing I want to build" in a way that has been solved in many other parts of the stack.