4 ms·
csvkit makes displaying CSV in a terminal trivial and has all the tools to manipulate/filter data I've ever needed - https://csvkit.readthedocs.io/en/latest/ ht
by strunz 3y ago
csvkit makes displaying CSV in a terminal trivial and has all the tools to manipulate/filter data I've ever needed - https://csvkit.readthedocs.io/en/latest/ https://csvkit.readthedocs.io/en/latest/
I don't really get this project at all.
- vidarh 3y agoThe thing is, while I'll probably just stick with CSV too, I'm sympathetic to the intent, but given I expect it'll need tooling anyway I'm less sympathetic to them not picking the existing separator. I also think there are failed lessons here that reduces the incentive for switching. E.g. If you're going to improve on CSV, a key improvement would be to aim to make the format trivially splittable, because the lesson from CSV is that when a format looks this trivial people will assume they can just split on a fixed string or trivial regex, and so the more you can reduce the harm of that the better. As such, I'd avoid most of the escaping they show, especially for line endings, and just make RS '\n' the record separator, or possibly RS '\n'*. Optionally do the same for US. Require escaping LF immediately after RS/US, and only allow escaping RS, so unescaping can be done with a trivial fixed replace per field if you have a reason to assume your data might have leading linefeeds in fields - a lot of apps will get away with just ignoring that. Then parsing is reduced to something like `data.split(RS).map{|row| row.split(US).map{|col| col.gsub(ESCAPE,"\n") } }` (assuming RS, US, and ESCAPE are regexps that include the optional trailing linefeeds and escapes leading linefeeds respectively). Being able to copy a correct one-liner from Stackoverflow ought to avoid most of the problems with broken CSV/TSV parsing. I'm also not convinced adding GS, FS, ETB is a good idea, partly for that reason, partly because a lot of the tools people will want to load data into will not handle more than one set of records, and so you'll end up splitting files anyway, in which case I'd just use a proper archive format... Those characters feels like they're trying to do too much given they're "competing" primarily with CSV/TSV. Their spec also needs to talk about encoding, because unless I've missed something, they only talk about codepoints, and they're likely to e.g. get people splitting on the UTF8 sequence etc. This to me is another reason for using the ASCII values - they encode the same in ASCII based characters sets and UTF8, and so it feels likely to be more robust against the horrors of people doing naive split-based parsing.
- strunz 3y agoCSV isn't even restricted to comma as the separator. You can use any character you like (pipe | is a common one) and csvkit will happy still work with a simple CLI flag. Pretty much all Unix tools have a similar flag. I've always been able to find an ASCII character that my data doesn't use, though maybe there are exceptions I haven't hit.
- bdzr 3y agoI love csvkit, particularly csvstat. I just wish it were quicker on larger files. The types I deal with routinely take 5-20 minutes to run and those are usually the ones I want the csvstat output for the most.
- shawn_w 3y agoI've been tempted a few times to rewrite some csvkit utilities in a faster language than Python.