3 ms·
Obligatory DuckDB solution: > duckdb -s "COPY (SELECT url[20:] as url, date, count(*) as c FROM read_csv('data.csv', columns = { 'url': 'VARCHAR', 'date': 'DAT
by csjh 7mo ago
Obligatory DuckDB solution:
> duckdb -s "COPY (SELECT url[20:] as url, date, count(*) as c FROM read_csv('data.csv', columns = { 'url': 'VARCHAR', 'date': 'DATE' }) GROUP BY url, date) TO 'output.json' (ARRAY)"
Takes about 8 seconds on my M1 Macbook. JSON not in the right format, but that wouldn't dominate the execution time.
- cess11 7mo agoThis log in one of the PR:s claims a 5.4s running time on some Mac. https://github.com/tempestphp/100-million-row-challenge/pull/47/changes#diff-7d100751a9dca9704b546197b4cb78f716a7a142765e75787ce4fcd97758ae52 https://github.com/tempestphp/100-million-row-challenge/pull...