3 ms·
They answer their own question: >Most of us don’t use punch-cards anymore, but that ease of authorship remains one of CSV’s most attractive qualities. CSVs can
by dbreunig 5y ago
They answer their own question:
>Most of us don’t use punch-cards anymore, but that ease of authorship remains one of CSV’s most attractive qualities. CSVs can be read and written by just about anything, even if that thing doesn’t know about the CSV format itself.
Yes, we keep CSVs. If you care about metadata and incredibly strict spec compliance, then yes: avro, parquet, json, whatever. But most CSV usage is small data, where the ease of usage, creation, and evaluation wins.
One of the problems with CSVs he cites is a great reason why I like CSVs:
>CSVs often begin life as exported spreadsheets or table dumps from legacy databases, and often end life as a pile of undifferentiated files in a data lake, awaiting the restoration of their precious metadata so they can be organized and mined for insights.
A benefit of a CSV is a skilled or unskilled operator can evaluate these piles of aging data. Parquet? Even SQLite? Not so much.
For small-to-medium sized datasets, CSV is great and accessible to a wider user base. If you're relying on CSV to preserve structure and meta, then meh.