5 ms·
I am curious about the motivation you choose Clickhouse over Apache Pinot, and Apache Druid? It could be helpful for other folks when choosing the OLAP db from
by metahunter 3y ago
I am curious about the motivation you choose Clickhouse over Apache Pinot, and Apache Druid? It could be helpful for other folks when choosing the OLAP db from one of them.
- vadman97 3y agoFor us, a significant reason was the ClickHouse cloud-hosted offering, rather than having to manage a cluster ourselves. Their use of S3 as the backing storage medium means that large-scale data retention is quite affordable. A good comparison we've referenced: https://leventov.medium.com/comparison-of-the-open-source-olap-systems-for-big-data-clickhouse-druid-and-pinot-8e042a5ed1c7 https://leventov.medium.com/comparison-of-the-open-source-ol...
- deleted 3y ago[deleted]
- dimitrios1 3y agoWhen I was highly engaged with Imply (Druid) a few years ago, S3 was also used as a backing storage. Is this not the case anymore?
- grumblestumble 3y agoFor reference, Apache Druid has an equivalent in Imply Polaris, and Apache Pinot has an equivalent in Startree. I can't speak for Startree, but Polaris similarly uses S3 for backing.
- metahunter 3y agoI think both Pinot and Druid nowadays offer cloud-hosted solutions. Maybe you started early that only ClickHouse had that offering. Is cloud hosting the only reason you guys choose Clickhouse? I am also wondering is it possible to let users choose the data source?
- berkle4455 3y agoClickhouse is fast and doesn’t have absurd architectural complexity.
- podoman 3y ago+1
- frankjr 3y agoNot having to deal with a JVM is a major plus tbh.
- douglasisshiny 3y agoI've seen so many variations of this comment on HN and I'm still not sure why not having to deal with the JVM is a major plus.
- preseinger 3y agobasically, the jvm is technically sophisticated but operationally complicated it sucks to use many people believe otherwise, but those people have rich jvm experience, which is not easy to get
- wpietri 3y agoI'm perfectly fine with JVMs, but at a guess, some of it is the usual snobbery for anything strange. But some of it is due to associating JVMs with enterprise nightmares. And some is that JVM tuning is a bit of a dark art. I've made some very good money going in and turning JVM knobs that others were afraid to touch. (The secret, by the way, is to hack together some decent load simulation and then measure not just median numbers but things like 99th percentile latency.)
- vetrom 3y agoJVM runtimes have a relatively high startup cost, are not often good 'citizens' in an instance running multiple types of software, and the build processes for a lot of JVM deliverables is an ungodly mess. Many of those bells and whistles are near-necessary in the enterprise world, but you have the accumulated mass of 'red zones' and developmental landmines in that ecosystem that can quickly turn you off it as a whole if you want to understand the whole system.
- douglasisshiny 3y agoI still don't understand some of this -- I developed in Java for 5+ years. >JVM runtimes have a relatively high startup cost I think many people are okay with that when developing server software that's going to run weeks at a time. It can get a bit annoying with trying to rapidly iterate. And I think things are changing pretty quickly with AOT builds and general improvements. >and the build processes for a lot of JVM deliverables is an ungodly mess. I recall using "mvn package." That's it. This was on two different systems that served a good bit of traffic and weren't simple trivial projects.
- drowsspa 3y agoDruid has like 9 different node types and inherits the whole Hadoop configuration mess and complexity
- grumblestumble 3y ago3, and there's absolutely no need for hadoop, particularly with MSQ
- tnolet 3y agoAnecdata: tried out Druid and Clickhouse for my SaaS. Couldn’t get Druid working. CH ran in 2 minutes.
- tnolet 3y agoSlight hijack. I / we went through a very similar tech selection process for timeseries metrics (not logging) ~1.5 years ago. We looked at Druid, ElasticSearch, TimeScale and a bunch of others. Main takeaways were: the SQL flavor and its aggregations in CH are amazing. Running on a single node for dev laptops is trivial. It’s crazy fast with almost zero tuning. It does not surprise me at all the CH is powering new products and startups. Note: hosted CH did not exist yet. We are using Altinity to run our cluster.
- preseinger 3y agoif you can afford SQL then you're not really doing timeseries in any meaningful sense
- podoman 3y ago> Note: hosted CH did not exist yet. We are using Altinity to run our cluster. It exists now actually. We (highlight) are on hosted clickhouse, which went in GA a few months ago. https://clickhouse.com/cloud https://clickhouse.com/cloud
- hodgesrm 3y agoThanks for the shout out! "Altinity" in this case means Altinity.Cloud, which is a high-performance cloud ClickHouse. It's been around for over 2.5 years. Disclaimer: I work at Altinity.
- metahunter 3y agoInteresting, just found another post from yesterday about the comparison: https://news.ycombinator.com/item?id=35642522 https://news.ycombinator.com/item?id=35642522, though the comparison is coming from Pinot team.