6 ms·
I wonder what niche these tools fill. Maybe when you want to do some ad-hoc analysis and want every step in your shell history? Otherwise it seems more flexibl
by fyp 7y ago
I wonder what niche these tools fill. Maybe when you want to do some ad-hoc analysis and want every step in your shell history?
Otherwise it seems more flexible to just fire up a python interpreter and do it in like 3 lines of pandas (or with sqlite and .import like another commenter mentioned)
- prepend 7y agoI’ve never had to use joins, but I’ve used shell scripts in CI environments where I can’t install stuff. For example, I had a Jekyll site generated from a project in some Linux ruby container. I couldn’t install python (or this xsv tool) but still wanted to summarize some of the data for display in one of the pages. While I would certainly rather use python, usinf shell was easier than convincing the admin of that server to install python into the ruby image, or trying to learn ruby and however they do simple data counts. This tool is weird because if I’m going to download some tool that’s not in the distro, I’d rather use SQLite or try to figure it out with awk.
- vvoyer 7y ago+1 for python panda. Amazing tool to manipulate data and csv very easily. Lots of options!
- DasIch 7y agoFor me it's ad-hoc analysis on large CSV files. Large meaning well beyond what Excel would be capable of, often larger than fits into memory on my local machine (10s of GiB). Sometimes I also use xsv to just do a step of the analysis and dive deeper on some subset using pandas. In my experience both SQLite and Pandas aren't as fast as fast for large files. So they are not really good options. Pandas is especially bad because it uses a column oriented data structure internally so reading from or writing to CSV is incredibly slow in Pandas. If you can use parquet that's not a problem but unfortunately parquet is not nearly is ubiquitous as csv :(
- nooorofe 7y agoIf Pandas is slow, than you can use Spark. For such big files laptop is not an option anyway. SQLite can be fast if you index your data (but I've worked with files < 10G). Nowadays I am just uploading CSV to some cloud database and work with data there.
- makapuf 7y ago> For such big files laptop is not an option anyway Too big for excel is not big data, and my laptop can load this 10G in RAM (not that it necessarily need all of it) so why not if the data is here and the laptop on your lap ?
- Scarbutt 7y agoThe same niche grep, sed, cut, sort, etc.. fill but specifically for csv and, using them together (xsv with other unix tools) can get you quickly very far on some tasks.
- SkyPuncher 7y agoI had to do this type of work frequently at an old job. We'd built short-run software systems for clients. Often this resulted in one-off data loads or reporting metrics. It was almost always easier just to knock something out with command line tools than it was to write, test, and deploy code. The deployment process itself could take longer than actually processing the file.
- arjie 7y agoThe shell is a very powerful REPL for incremental development, so you can often do things more rapidly by looking with `xsv headers`, then joining with `xsv join | xsv select | head | xsv table` than using python or sqlite. I definitely used to use `sqlite` for every CSV before and I've shifted to using `xsv`. If I want to do any heavy lifting I'm going to pop it into PostgreSQL anyway. Crucial feature is that it's easy to pop some `xsv` pipeline in the middle of your `for i in ...` or `find . -print0 | xargs -0`.