5 ms·
This looks great! Please consider removing any implicit network calls like the initial "Checking GitHub for updates...". This itself will prevent people from a
by snidane 3y ago
This looks great!
Please consider removing any implicit network calls like the initial "Checking GitHub for updates...". This itself will prevent people from adoption or even trying it any further. This is similar to gnu parallel's --citation, which, albeit a small thing - will scare many people off.
Consider adding pivot and unpivot operations. Mlr gets it quite right with syntax, but is unusable since it doesn't work in streaming mode and tries to load everything into memory, despite claiming otherwise.
Consider adding basic summing command. Sum is the most common data operation, which could warrant its own special optimized command, instead offloading this to external math processor like lua or python. Even better if this had a group by (-by) and window by (-over) capability. Eg. 'qsv sum col1,col2 -by col3,col4'. Brimdata's zq utility is the only one I know that does this quite right, but is quite clunky to use.
Consider adding a laminate command. Essentially adding a new column with a constant. This probably could be achieved by a join with a file with a single row, but why not make this common operation easier to use.
Consider the option to concatenate csv files with mismatched headers. cat rows or cat columns complains about the mismatch. One of the most common problems with handling csvs is schema evolution. I and many others would appreciate if we could merge similar csvs together easily.
Conversions to and from other standard formats would be appreciated (parquet, ion, fixed width lenghts, avro, etc.). Othe compression formats as well - especially zstd.
It would be nice if the tool enabled embedding outputs of external commands easily. Lua and python builtin support is nice, but probably not sufficient. i'd like to be able to run a jq command on a single column and merge it back as another for example.
Inspiration:
- csvquote: https://news.ycombinator.com/item?id=31351393
- teip: https://github.com/greymd/teip
- quasarj 3y agoWait, who is scared off by parallel's --citation?
- fbdab103 3y agoI refuse to use parallel due to that obnoxiousness. At minimum, it is not installed by default, so it is already a negative to just using xargs. That it then puts that barrier in my way makes it an easy tool to skip.
- quasarj 3y agoI just don't understand what barrier you are talking about. I just checked, it doesn't even whine at you when you use it, the help just notes that you should cite it if you publish a paper where you used it. And... anyone publishing papers knows about citation requirements lol. Anyone else can ignore it. What is this barrier?
- dima55 3y agoIn addition to being annoying, it raises questions about whether it is free software or not. Some people care a whole lot about that. And some people have higher standards about being nagged. And lots and lots of time was spent discussing solutions, for instance: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=915541 https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=915541
- quasarj 3y agoAh, I see, they have changed it (or possibly the version on my system has had the --will-cite patched out, as discussed in this bug). Okay, I accept your argument about Free Software. However, I find it interesting that it's a GNU project... they are generally the most hardline Free Software people.
- fbdab103 3y agoTo slippery slope this, what happens if more tools start adopting this behavior? Curl now asks you to buy Daniel Stenberg a coffee on each use. Wget asks you to support Ukraine. Caddy wants you to invest in their startup. Each of which may come with their own `--ignore-annoyance-flag` I need to learn. The best I can do is vote with my feet. I also do not care for the citation requirement. I utilize tons of tools in my work which go unstated. I do not feel the need to cite Linux, DNS, htop, Make, Diet Coke, my Kinesis keyboard, etc. Sadly, reliable plumbing gets no respect. Especially for a tool which is more or less interchangeable with some shell scripting. Unless I am trying to shore up the references list, I am going to cite directly relevant work. At some point, you no longer need to note that your work was powered by electricity.
- sweetgiorni 3y agoI find it incredibly obnoxious and I refuse to use parallel because of it. To me, it violates the spirit of free software and tarnishes the GNU project. As someone who has released my source to the public for free, I couldn't fathom adding such a flag. Bonus SO post to enhance your fury: https://stackoverflow.com/questions/61762189/installing-gnu-parallel-how-to-enter-will-cite-from-docker-build https://stackoverflow.com/questions/61762189/installing-gnu-...
- dima55 3y agoYou can get quite far by piping to other tools and/or using DSLs. pivoting can almost certainly be done by the luau support in qsv (or `vnl-filter`, for instance). Summing and grouping is something that `datamash` does well (or qsv luau probably, or `vnl-filter --eval`). Adding a column once again can be done with luau or `vnl-filter`. Would you be more likely to use this tool if it had even more stuff in it requiring reading even more documentation? That's a genuine question.
- ezequiel-garzon 3y agoI know this is just one thing out of many, but sum is included in stats.
- jqnatividad 3y agoThanks for the detailed feedback @snidane! As maintainer of qsv, here's my reply: - Given qsv's rapid release cycle (173 releases over three years), the auto-update check is essential at the moment. Once we reach 1.0, I'll turn it off. For now, given your feedback, I've only made it check 10% of the time. - Pivot is in the backlog and I'll be sure to add unpivot when I implement it. (https://github.com/jqnatividad/qsv/issues/799 https://github.com/jqnatividad/qsv/issues/799) - I'll add a dedicated summing command with the group by (-by) and window by (-over) capability (https://github.com/jqnatividad/qsv/issues/1514 https://github.com/jqnatividad/qsv/issues/1514). Do note that `stats` has basic sum as @ezequiel-garzon pointed out. - With the `enum` command, qsv can achieve what you proposed with `laminate`. E.g. qsv enum --new-column newcol --constant newconstant mydata.csv --output laminated-data.csv - With the cat rowskey command, qsv can already concatenate files with mismatched headers. - other file formats. qsv supports parquet, csv, tsv, excel, ods, datapackage, sqlite and more (see https://github.com/jqnatividad/qsv/tree/master#file-formats https://github.com/jqnatividad/qsv/tree/master#file-formats). Fixed-format though is not supported yet and quite interesting, and have added it to the backlog (https://github.com/jqnatividad/qsv/issues/1515 https://github.com/jqnatividad/qsv/issues/1515) - as to "enable embedding outputs of commands", qsv is composable by design, so you can use standard stdin/stdout redirection/piping techniques to have it work with other CLI tools like jq, awk, etc. Finally, just released v0.120.0 that already incorporates the less aggressive self-update check. https://github.com/jqnatividad/qsv/releases/tag/0.120.0 https://github.com/jqnatividad/qsv/releases/tag/0.120.0