3 ms·
As someone who likes modern formats like parquet, when in doubt, I end up using CSV or JSONL (newline-delimited JSON). Mainly because they are plain-text (fast
by polyrand 2y ago
As someone who likes modern formats like parquet, when in doubt, I end up using CSV or JSONL (newline-delimited JSON). Mainly because they are plain-text (fast to find things with just `grep`) and can be streamed.
Most features listed in the document are also shared by JSONL, which is my favourite format. It compresses really well with gzip or zstd. Compression removes some plain-text advantages, but ripgrep can search compressed files too. Otherwise, you can:
zcat data.jsonl.gz | grep ...
Another advantage of JSONL is that it's easier to chunk into smaller files.
- sitkack 2y agoI switched to JSONL over a decade ago and I would recommend everyone else to also have switched then. This whole thread is an uninformed rehash of bad ideas.
- theLiminator 2y agoI think that might make sense ingest side, but that's very expensive to deal with if you're doing anything remotely large. I think sinking into something like delta-lake or iceberg probably makes sense at scale. But yeah, I definitely agree that CSV is not great.
- sitkack 2y agoJSONL as a replacement for CSV, you shouldn't be using CSV as format for long term storage or querying, it has so many downsides and nearly zero upsides. JSONL when compressed with zstd, most of "expensive if large" disappears as well. Generating and consuming JSONL can easily be in the GB/s range.
- theLiminator 2y agoI mean on the querying side. Parquet's ability to skip rowgroups and even pages, and paired with iceberg or delta can make the difference between being able to run your queries at all versus needing to scale up dramatically.
- sitkack 2y agoTotally agree. I am saying JSONL is a lower bound format, if you can use something better you should. Data interchange, archiving, transmission, etc. It shouldn't be repeatedly queried. Parquet, Arrow, sqlite, etc are all better formats.
- dsp_person 2y agoToo bad xz/lzma isn't supported in these formats. I often get pretty big improvements in compression ratio. It's slower, but it can be parallelized too.
- cbsmith 2y agoYou can cat Parquet and other formats into grep just as easily.