3 ms·
"Unsaid"? I have the opposite perspective: This is repeated ad nauseam. Every time a binary format or any other complex optimization is mentioned, the performan
by rcoveson 4y ago
"Unsaid"? I have the opposite perspective: This is repeated ad nauseam. Every time a binary format or any other complex optimization is mentioned, the performance-doesn't-matter people have to come out of the woodwork and bring up the fact that it "literally doesn't matter" to "$INVENTED_HIGH_PERCENTAGE of use cases".
I wonder, do people on the Caterpillar forum have this problem, where people just show up to ask, "Yeah but can you take it on the freeway? Because I can take my Tacoma on the freeway."
- agiacalone 4y agoI’m confused. Are you advocating for complexity regardless of need?
- rcoveson 4y agoHow could anybody advocate for "complexity regardless of need"? What would that even look like? Maybe you were being sarcastic, I'm not sure. But I hope it's obvious that advocating for complexity at all is not the same thing as advocating for complexity regardless of need.
- FridgeSeal 4y agoIf you are under the impression that CSV’s-and-friends are “less complex” because they’re text, I would like to assure you that any “simplicity” in the existing text formats is a heinous lie, and the complexity gets pushed into the code of the consuming layer. I use parquet for small stuff now, because it’s standardised and so much nicer and more reliable to use. It’s faster. It ships a schema. There’s actual data type support. There’s even partitioning support if you need it. It’s quite literally better in every useful way. Long winded way of saying: the complexity has always been there. CSV’s act as if it’s not there, parquet acknowledges it and gives you the tools to deal with it.
- syntheweave 4y agoThe angle by which CSV is simple is in the flexibility of input...that is, it's a very decent authoring format, when sent into a system that knows how to import and validate "that kind of CSV" and treat it the way all user interfaces should treat input - with an eye towards telling the user where they screwed up. Once you've got the data into a processing pipeline you do need additional semantic structure to not lose integrity, which I think Parquet is fine for, but any SQL schema will also do.
- cout 4y agoParquet is not better than csv in every useful way. CSV, being row-oriented, makes it much easier to append data than parquet, which is column-oriented. CSV is also supported by decades of software, much of which either predates parquet or does not yet support parquet. CSV, being a newline-delimited text-based format, is much better suited for inclusion in a git repository than parquet. I use parquet wherever I can, but I still reach for csv whenever it makes sense.
- heavenlyblue 4y ago> CSV is also supported by decades of software, much of which either predates parquet or does not yet support parquet. CSV can not be supported by any software because it's not really a format. If you don't understand what I am saying by this, please write down an algorithm for figuring out which type of quoting is used in a given CSV file. Also please explain to me how you would deal with new lines in the text data.
- yardstick 4y agoI’ve come here to ask why CSV is being used? CSV is easy for humans to read, and works in almost any tool. It’s Excel friendly (yes Excel has some terrible CSV/data mangling practices). End of the day, if the format needs to be used by arbitrary third parties who maybe/probably have no technical experience beyond excel, then CSV is the best option. If human readable and Excel support are not required, then by all means Parquet is the winner. TLDR: The best format depends on how you need to use the data.