4 ms·
The author had no idea what they were doing with kdb. They even admit that they couldn't be bothered to modify their ingestion scripts to not partition their d
by ducreux 8y ago
The author had no idea what they were doing with kdb.
They even admit that they couldn't be bothered to modify their ingestion scripts to not partition their data.
- gricardo99 8y agoYup. >The pickup_ntaname column is stored as varchars by ClickHouse and kdb+, and as dictionary encoded single byte values by LocustDB. But it would be trivial to convert to enum/sym type in kdb+. It's silly to query and group by strings.
- inteleng 8y agoThis is one of the frustrating parts of database software. Where can information about the ways to optimize these parameters be found (outside of random posts scattered around StackExchange)?
- gricardo99 8y agoso true. Either you have lots of experience working with specific databases to setup/optimize queries, and you know what works best through personal blood, sweat and tears, and/or you have intimate knowledge of the inner workings of the database architecture/implementation and know the theoretical best approach to structure your schema/queries. But even then, hardware/networking performance and tuning can throw a wrench in the most seasoned/knowledgeable approaches. Users can further bring otherwise solid setups to a grinding halt with unanticipated use-cases. The only hope when you hit these inevitable road-blocks is that you're working for someone that appreciates the difficulty of the problem.
- frankmcsherry 8y agoThis seems like a very unfair reading of what the author actually wrote: > One note about the results for kdb+: The ingestion scripts I used for kdb+ partition/index the data on the year and passenger_count columns. This may give it a somewhat unfair advantage over ClickHouse and LocustDB on all queries that group or filter on these columns (queries 2, 3, 4, 5 and 7). I was going to figure out how to remove that partitioning and report those results as well, but didn’t manage before my self-imposed deadline.
- anonu 8y agoCompletely moot point though as he demonstrates that even when kdb+ is advantaged by having data be indexed, LocustDB is still faster in 4 of the 5 queries he runs... So yeah, maybe the guy has no idea what to do with kdb - but ultimately having a fast, free and open-source database & query language beats a fast and expensive piece of commercial software.
- geocar 8y agoThe first red flag is that Mark's benchmarks look very different for kdb, even though his ClickHouse times are similar to Clemens. Looking over the queries, he made some... interesting changes that have him benchmarking oranges to apples.