4 ms·
That is what tipped me over the edge for this little rant. I've talked to a lot of folks who don't have a background in the data ecosystem and want to do fairly
by memset 2y ago
That is what tipped me over the edge for this little rant. I've talked to a lot of folks who don't have a background in the data ecosystem and want to do fairly mundane things ("copy data out of a parquet file into a regular text file") and just hit this wall where they have to learn to deal with edge case after edge case.
The person in that comment, they didn't choose to get this dump of data in this format, but now they have to contend with it, and it sucks that he - someone who's a very capable dev - had to spend to much time on this problem.
- wenc 2y agoI think there’s a difference between saying “ok I don’t know how to do this, what’s an easier way?” Rather than going on a rant that “bunch of parquet files sucks!” When it’s clearly untrue and all that was needed was knowledge of the tools. Reading a ton of data from an unfamiliar format (but one that is specialized toward solving a problem) is not mundane. It’s not meant to be easy. Parquet feels like it should be easy and many tools have been built to make it easy (how far we’ve come since the Java parquet reader!) but it still requires knowing what to do. Not knowing at first doesn’t mean the right thing to do is to fall back to the wrong but simple solution of CSV. I’ve seen so many people do that in my company’s data pipelines and it causes so many downstream issues. This is the wrong message to send that will incur so much cost for so many companies if believed. It’s cool that you’re making tools to make things easier. you can do that without ranting about parquet which just seems misdirected.
- memset 2y agoThe article isn't particularly about Parquet or S3. They are both great technologies and I use them all the time. I mentioned CSV because it is the only format supported by Postgres for ingest (besides their binary protocol) but that is only an instance of the general problem of interoperability. I'm not sure I agree that "reading a ton of data is not meant to be easy." I'd say it is not easy today, and there is a constellation of tools and programming techniques that can let you perform the task, but it is often a distraction from the end result you want to achieve. The essence of my rant is that there are many steps distracting steps between "a data dump" and "the thing I want to build" in a way that has been solved in many other parts of the stack.