3 ms·
Interesting, why it would take so long ? I'm genuinely curious so that we can make it better (Disclaimer:I work at ClickHouse) Here's a example of a query agai
by ryadh 4y ago
Interesting, why it would take so long ? I'm genuinely curious so that we can make it better (Disclaimer:I work at ClickHouse)
Here's a example of a query against a Parquet file you can run on your laptop:
```
./clickhouse local -q "SELECT
town,
avg(price) AS avg_price
FROM s3('https://datasets-documentation.s3.eu-west-3.amazonaws.com/house_parquet/house_0.parquet https://datasets-documentation.s3.eu-west-3.amazonaws.com/ho...')
GROUP BY town
ORDER BY avg_price DESC
LIMIT 10"
```
From:
https://clickhouse.com/docs/knowledgebase/parquet-to-csv-json#accessing-the-data-using-a-table-function https://clickhouse.com/docs/knowledgebase/parquet-to-csv-jso...
- dmw_ng 4y agoBig fan of ClickHouse, but in the case of S3 there doesn't seem to be much point fighting city hall. I was the author of the original clickhouse-local-in-Lambda ticket FWIW. Even if that approach worked well, it'd still only be <50% cheaper than Athena's fractions-of-a-cent for queries on efficiently partitioned data, and still there is setup work involved. I think CH-local-in-Lambda might still easily compete on cost with Athena for raw CSV queries though
- ryadh 4y agoInteresting, thanks for sharing this feedback! I didn't realise intially that it was about running clickhouse-local in Lambdas. Fyi, we recently improved the Parquet support in 23.2 https://github.com/ClickHouse/ClickHouse/pull/45878 https://github.com/ClickHouse/ClickHouse/pull/45878 Also, we still have optimizations for reading Parquet from S3 coming so that might improve