8 ms·
Maybe we need a reminder every now and then that text-files are quite capable and except from being faster to parse in a lot of situation has the benefits of be
by Puts 5y ago
Maybe we need a reminder every now and then that text-files are quite capable and except from being faster to parse in a lot of situation has the benefits of being "future safe" and easy to backup and compress as well.
- spullara 5y agoI mean, it isn't like Hadoop wasn't used to parse text files. Also, it is all fun and games until you need a join.
- kragen 5y agocomm(1) and join(1) can do joins on text files. Make sure they're sorted in the same locale you're joining them in; LANG=C tends to be the fastest.
- IanCal 5y agoThe common text file for lots of this, CSVs, are absolutely awful. They're fine until they totally aren't, they were just the best option for a lot of use cases. I'd argue that's now been entirely replaced with parquet, significantly faster and broad support, proper types and more.
- stingraycharles 5y agoI’d argue that for as long as Excel doesn’t support Parquet files, we haven’t seen the last of CSVs for a long time. Parquet is great, but it’s simply nowhere near as ubiquitous as CSV.
- gehen88 5y agoI hadn't even heard of Parquet until now, and I'm sure this goes for lots of developers who don't do much data engineering.
- jsjohnst 5y agohttps://xkcd.com/1053/ https://xkcd.com/1053/ What’s the Parquet equivalent of going to the store to buy Mentos and Diet Coke now?
- tharne 5y ago> I’d argue that for as long as Excel doesn’t support Parquet files, we haven’t seen the last of CSVs for a long time. It's easy to forget just how much analytic "stuff" Excel still powers.
- bell-cot 5y agoIf you don't already have parquet deployed, there's a wide gulf (in skill set, overhead, etc.) between CSV and parquet. If you don't mind being old-school, the data is ASCII text, and you're tired of some of CSV's little issue, then ASCII has the FS, GS, RS, and US control characters - specifically intended for such uses. Micro conceptual overhead, none of the CSV issues which screw up many *nix text-handling programs and little script files, and a decent modern filesystem can handle the compression separately.
- kenoph 5y agoTbh Unix programs don't handle non-ASCII text very well, in my experience.
- derriz 5y agoYes - parquet tooling is non-existent compared to CSV particularly on the command line. And the cross-language/platform support is a mess - good luck reading Pandas generated parquet on .NET or in a (non-spark) JVM environment. There are many reasons why CSV is flawed for the purposes of storing tabular data (e.g. loss of column type information) but the alternatives are just so unergonomic that CSV remains a viable choice in many situations.
- manigandham 5y agoAlternatives are available like Avro
- IanCal 5y agoFor many cases of passing data between systems, I'd say the gulf is dramatically lower than it has been until even fairly recently. Support for many standard packages is just there, swapping out "read_csv" for "read_parquet" if you're a pandas shop may even be enough. More and more tools read these directly, and it opens up a whole load of better options for processing data. > If you don't mind being old-school, the data is ASCII text, and you're tired of some of CSV's little issue, then ASCII has the FS, GS, RS, and US control characters I have used those before, and yet I still had those characters appear in data. The only places I'd ever seen them were in the wiki page and in customer delivered data. Absolute pain to dig through and remove. On top of that "if your data is ASCII" is something I'd be nervous about for many use cases even if it is right now. Beyond that, then you need everyone to swap out their parsing to use those characters. CSV is fine until it totally blows up in your face. All it takes is one "oh it's fine we'll use awk" stage somewhere or a CSV parser that isn't good enough and one person to put a newline where nobody had expected it before.
- dhx 5y agoColumn-oriented formats such as Parquet can be awful too. For example, if you have five columns of numbers you want to multiply together, Parquet is going to be the worst option available because of the extreme cache thrashing that the CPU will encounter. Structure packing[1] and consideration of locality of reference[2] would need to be applied for high performance applications where a computer scientist has considered the algorithm needing to be implemented and the most efficient data format that the source data would need to be provided in. [1] http://www.catb.org/esr/structure-packing/ http://www.catb.org/esr/structure-packing/ [2] https://en.wikipedia.org/wiki/Locality_of_reference https://en.wikipedia.org/wiki/Locality_of_reference