4 ms·
Personally, I use xsv and it’s been tremendously helpful, especially when working with larger files. https://github.com/BurntSushi/xsv https://github.com/BurntS
by sklarsa 3y ago
Personally, I use xsv and it’s been tremendously helpful, especially when working with larger files. https://github.com/BurntSushi/xsv https://github.com/BurntSushi/xsv
- elmolino89 3y agoxsv is great for a quick sanity checks (i.e. number of columns, unique values counts in a given column) but for a more serious tasks/giant files I switch to either polars or duckdb converting CSV/TSV files to parquet or parquet data sets. By giant I mean 25G gzipped files with >10^9 rows like these VCFs: https://ftp.ncbi.nlm.nih.gov/snp/latest_release/VCF/ https://ftp.ncbi.nlm.nih.gov/snp/latest_release/VCF/
- rzmk 3y agoMaybe try the to [1] and sqlp [2] commands from the qsv fork. From the README: sqlp: Run blazing-fast Polars SQL queries against several CSVs - converting queries to fast LazyFrame expressions, processing larger than memory CSV files. to: Convert CSV files to PostgreSQL, SQLite, XLSX, Parquet and Data Package. [1] https://github.com/jqnatividad/qsv/blob/master/src/cmd/sqlp.rs#L2 https://github.com/jqnatividad/qsv/blob/master/src/cmd/sqlp.... [2] https://github.com/jqnatividad/qsv/blob/master/src/cmd/to.rs#L2 https://github.com/jqnatividad/qsv/blob/master/src/cmd/to.rs...
- keybored 3y agoI shared the qsv fork [1] yesterday which is more active. xsv is more lean while qsv tries to support every action that you might want to perform on CSV files.[2] [1] https://github.com/jqnatividad/qsv https://github.com/jqnatividad/qsv [2] https://github.com/jqnatividad/qsv/discussions/290#discussion-4075202 https://github.com/jqnatividad/qsv/discussions/290#discussio...