14 ms·
ClickHouse as an alternative to Elasticsearch for log storage and analysis
- Bluestein 6y agoRead that as "Clubhouse" on first glance ... ... and was quite intrigued :) PS: Tough crowd tonite ... It would appear local crowd does not take well to some tongue-in-cheek. Besides, it actually happened as described.-
- Jennifer9918910 6y agoI'm available for sex, Please just follow my link https://rebrand.ly/sexygirls007 https://rebrand.ly/sexygirls007
- wiradikusuma 6y agoAnyone know more lightweight alternative to (ELK) Elastic Stack? I found https://vector.dev https://vector.dev but it seems to be only the "L" part.
- webo 6y agoHave you looked into Google Cloud Logging (Stackdriver)? It's the most affordable and decent-enough solution we've found. The only issue is querying can be slow on large volumes.
- ehnto 6y agoI found Lucene's base library really easy to use without the configuration/infrastructure overhead of Elasticsearch, but haven't experienced it at scale: https://lucene.apache.org/ https://lucene.apache.org/
- rzzzt 6y agoSolr is the equivalent-ish of ES, if you are looking for a search server instead of a library that can be embedded: https://lucene.apache.org/solr/ https://lucene.apache.org/solr/
- Xylakant 6y agoPromtail/Loki https://github.com/grafana/loki https://github.com/grafana/loki is an alternative to elk, but while it seems more lightweight, it definitely is less featureful. The integration with grafana/prometheus seems nice, but I've only toyed with it, not used in production.
- johnx123-up 6y agoWhat about https://github.com/meilisearch/MeiliSearch https://github.com/meilisearch/MeiliSearch ?
- wdb 6y agoI am looking into this. Do you have experience with it?
- johnx123-up 6y agoIt is good. I can't find any CDC for Postgres for the incremental sync. And so I had to use the bulk update/sync and that causes performance issues occasionally. Also, some Algolia features are not available yet https://github.com/meilisearch/instant-meilisearch/issues/215 https://github.com/meilisearch/instant-meilisearch/issues/21...
- FridgeSeal 6y agoClickHouse will happily replace the ElasticSearch bit, and there’s a few open source dashboards you could use as a kibana stand in: - Metabase (with ClickHouse plug-in) - Superset - Grafana
- wakatime 6y agoA related database using ideas from Clickhouse: https://github.com/VictoriaMetrics/VictoriaMetrics https://github.com/VictoriaMetrics/VictoriaMetrics
- wikibob 6y agoAre you familiar with VictoriaMetrics? Can you elaborate on how it is similar and dissimilar to Clickhouse? What specific techniques are the same?
- ekimekim 6y agoThe core storage engine borrows heavily from it - I'll attempt to summarize and apologies for any errors, it's been a while since I worked with VictoriaMetrics or ClickHouse. Basically data is stored in sorted "runs". Appending is cheap because you just create a new run. You have a background "merge" operation that coalesces runs into larger runs periodically, amortizing write costs. Reads are very efficient as long as you're doing range queries (very likely on a time-series database) as you need only linearly scan the portion of each run that contains your time range.
- sylvinus 6y agoClickHouse is incredible. It has also replaced a large, expensive and slow Elasticsearch cluster at Contentsquare. We are actually starting an internal team to improve it and upstream patches, email me if interested!
- wikibob 6y agoCan you share some more details? How many nodes on both? How much data ingested and stored? What’s the query load?
- AurimasJLT 6y agohttps://github.com/ClickHouse/clickhouse-presentations/blob/master/meetup30/contentsquare.pdf https://github.com/ClickHouse/clickhouse-presentations/blob/... and the presentation itself https://www.youtube.com/watch?v=lwYSYMwpJOU https://www.youtube.com/watch?v=lwYSYMwpJOU 300Elastic nodes vs 12ClickHouse nodes / 260TB/ lots of querries
- jetter 6y agoYep, do you guys have a writeup on this? Altinity actually mention Contentsquare case in their video, here: https://www.youtube.com/watch?t=2479&v=pZkKsfr8n3M&feature=youtu.be https://www.youtube.com/watch?t=2479&v=pZkKsfr8n3M&feature=y...
- harporoeder 6y agoI don't have any production experience running Clickhouse, but I have used it on a side project for an OLAP workload. Compared to Postgres Clickhouse was a couple orders of magnitude faster (for the query pattern), and it was pretty easy to setup a single node configuration compared to lots of the "big data" stuff. Clickhouse is really a game changer.
- dominotw 6y agowhat happens when your data doesn't fit in a single node anymore?
- deleted 6y ago[deleted]
- qaq 6y agoIt scales to crazy numbers as you add nodes in 2018 CloudFlare was ingesting 11M rows per second into their CH cluster
- aseipp 6y agoThere's replication in ClickHouse and you can just shove reads off to one of them if you'd like. From a backup/safety standpoint that's important, but I think there are other options besides just replicas, of course. From an operations standpoint, however, ClickHouse is ridiculously efficient at what it does. You can store tens of billions, probably trillions of records on a single node machine. You can query at tens of billions of rows a second, etc, all with SQL. (The only competitor I know of in the same class is MemSQL.) So another thing to keep in mind is you'll be able to go much further with a single node using ClickHouse than the alternatives. For OLAP style workloads, it's well worth investigating.
- mekster 6y agoI don't know why people still call replication a backup.
- dominotw 6y agoHow does clickhouse compare to druid, pinot, rockset (commercial), memsql (commercial). I know clickhouse is easier to deploy. But from user's perspective is clickhouse superior to the others?
- jetter 6y agoI've mentioned Pinot and Druid briefly in 2018 writeup: https://pixeljets.com/blog/clickhouse-as-a-replacement-for-elk-big-query-and-timescaledb/ https://pixeljets.com/blog/clickhouse-as-a-replacement-for-e... (see "Compete with Pinot and Druid" )
- caust1c 6y agoWhen Cloudflare was considering clickhouse, we did estimates on just the hardware cost and it was well over 10x what clickhouse was based on druids given numbers on events processed per compute unit.
- advaita 6y agoAre you saying druid hardware costs were coming out to be 10x of clickhouse hardware costs? Caveat: English is not my first language so might have missed your point in translation. :)
- caust1c 6y agoThat's right, yeah. We would have had to buy 10x the servers in order to support the same workload that clickhouse could.
- AurimasJLT 6y agoYeah Druid got blown away by ClickHouse at eBay Druid 700+ servers versus 2 region fully replicated ClickHouse system of 40 nodes. https://tech.ebayinc.com/engineering/ou-online-analytical-processing/ https://tech.ebayinc.com/engineering/ou-online-analytical-pr... and a webinar they did with us at Altinity https://www.youtube.com/watch?v=KI0AqpmcSOk&t=20s https://www.youtube.com/watch?v=KI0AqpmcSOk&t=20s
- 6y ago
- BoorishBears 6y agoMy biggest problem with Elasticsearch is how easy it is to get data in there and think everything is just fine... until it falls flat on its face the moment you hit some random use case that, according to Murphy's law, will also be a very important one. I wish Elasticsearch were maybe a little more opinionated in its defaults. In some ways Clickhouse feels like they filled the gap not having opinionated defaults created. My usage is from a few years back so maybe things have improved
- djxfade 6y agoWould you care to elaborate on what happened in your case. My company is using ElasticSearch extensively, and it is mission critical for us. I fear something might happen one day
- BoorishBears 6y agoIt's been a few years so the details are fuzzy, but iirc it was just simple things like index sizing, managing shard count, some JVM tuning, certain aggregated fields blowing up once we had more data in our instances... We also got our data very out of order. We had embedded devices logging analytics that would phone home very infrequently, think months between check-ins I forget why but that became a big issue at some point, bringing the instance to its knees when a few devices started to phone-in covering large periods of time. ES just has a ton of knobs, I imagine if its been important to you, you have people specializing in keeping it running, which is great... but the amount of complexity there that is specific to ES is really high. It's not like there's no such thing as a Postgres expert for example, but you don't need to hire a Postgres wizard until you're pretty far in the weeds. But I feel like you should have an ES wizard to use ES, which is a little unfortunate
- scomp 6y agoI work for a company that uses it as a mission critical product. It provides the search in our SAAS. I'd say it mysteriously fails 3-4 times a year. But no one wants to invest or encourage anyone to look at it proactively or even re-actively. We've had JVM issues and indexers spiral out of control using 99%CPU only to stop at random hours later. It's definitely a product to learn before it fails.
- moralestapia 6y agoSorry to hijack the thread but can anyone suggest alternatives to the 'search' side of Elasticsearch? I haven't been following the topic and there's probably new and interesting developments like ClickHouse is for logging.
- ericcholis 6y agohttps://typesense.org/ https://typesense.org/ comes to mind. Has support for Algolia's instantsearch.js as well.
- drewda 6y agoNot sure if it's "new" but it's always interesting: Postgres offers a fine set of full-text search functionality, with the advantage of being able to also use the other ways in which Postgres shines: document-like data (storing JSON in columns), horizontal replication, PostGIS, and so on.
- stanmancan 6y agoSomeone recommended Meilisearch the other day. I've been playing around with it and it's pretty great so far. Still early in the project right now.
- abrookins 6y agoYou can use Redis for full-text search and some more SQL-like queries, including aggregations, with RediSearch: https://oss.redislabs.com/redisearch/ https://oss.redislabs.com/redisearch/
- jetter 6y agohttps://github.com/meilisearch/MeiliSearch https://github.com/meilisearch/MeiliSearch gets a lot of traction recently. There are also Sphinx and its fork https://manticoresearch.com/ https://manticoresearch.com/ - very lightweight and fast.
- tgtweak 6y agoI immediately thought of Sphinx when I saw MeiliSearch... it's uncanny how the use case and implementation semantics haven't changed much in 15 years. The beauty of pointing it to your mysql tables and getting fulltext-via-api on the other side was quite nice.
- js2 6y agoSentry.io is using ClickHouse for search, with an API they built on top of it to make it easier to transition if need be. They blogged about it at the time they adopted it: https://blog.sentry.io/2019/05/16/introducing-snuba-sentrys-new-search-infrastructure https://blog.sentry.io/2019/05/16/introducing-snuba-sentrys-...
- mekster 6y agoI like sentry but it's the only app that I know uses 25 or so containers to stich itself together to get running which seemed insane, not to mention the usage of much ram.
- moralsupply 6y agoI'm happy that more people are "discovering" ClickHouse. ClickHouse is an outstanding product, with great capabilities that serve a wide array of big data use cases. It's simple to deploy, simple to operate, simple to ingest large amounts of data, simple to scale, and simple to query. We've been using ClickHouse to handle 100's of TB of data for workloads that require ranking on multi-dimensional timeseries aggregations, and we can resolve most complex queries in less than 500ms under load.
- dabeeeenster 6y agoI've been recording a podcast with Commercial Open Source company founders (Plug! https://www.flagsmith.com/podcast https://www.flagsmith.com/podcast) and have been surprised how often Clickhouse has come up. It is always referred to with glowing praise/couldn't have built our business without it etc etc etc.
- jabo 6y agoThanks for sharing your podcast! Just subscribed. To add yet another data point, we use Clickhouse as well for centralized logging for the SaaS version of our open source product, and can't imagine what we would have done without it.
- tgtweak 6y agoI think it's an unfair comparison, notably because: 1) Clickhouse is rigid-schema + append-only - you can't simply dump semi-structured data (csv/json/documents) into it and worry about schema (index definition) + querying later. The only clickhouse integration I've seen up close had a lot of "json" blobs in it as a workaround, which cannot be queried with the same ease as in ES. 2) Clickhouse scalability is not as simple/documented as elasticsearch. You can set up a 200-node ES cluster with a relatively simple helm config or readily-available cloudformation recipe. 3) Elastic is more than elasticsearch - kibana and the "on top of elasticsearch" featureset is pretty substantial. 4) Every language/platform under the sun (except powerbi... god damnit) has native + mature client drivers for elasticsearch, and you can fall back to bog-standard http calls for querying if you need/want. ClickHouse supports some very elementary SQL primitives ("ANSI") and even those have some gotchas and are far from drop-in. In this manner, I think that clickhouse is better compared as a self-hosted alternative to Aurora and other cloud-native scalable SQL databases, and less a replacement for elasticsearch. If you're using Elasticsearch for OLAP, you're probably better to ETL the semi-structured/raw data out of ES that you specifically wan to a more suitable database which is meant for that.
- ignoramous 6y ago> Elastic is more than elasticsearch... Grafana Labs sponsored FOSS projects are probably adequate replacement for the Elasticsearch? https://grafana.com/oss/ https://grafana.com/oss/ > ...clickhouse is better compared as a self-hosted alternative to Aurora and other cloud-native scalable SQL databases Aurora would be likely be less better at this than RedShift or Snowflake.
- jetter 6y agoI address your concern from #1 in "2. Flexible schema - but strict when you need it" section - take a look at https://www.youtube.com/watch?v=pZkKsfr8n3M&feature=emb_title https://www.youtube.com/watch?v=pZkKsfr8n3M&feature=emb_titl... Regarding #2: Clickhouse scalability is not simple, but I think Elasticsearch scalability is not that simple, too, they just have it out of the box, while in Clickhouse you have to use Zookeeper for it. I agree that for 200 nodes ES may be a better choice, especially for full text search. For 5 nodes of 10 TB logs data I would choose Clickhouse. #3 is totally true. I mention it in "Cons" section - Kibana and ecosystem may be a deal breaker for a lot of people. #4. Clickhouse in 2021 has a pretty good support in all major languages. And it can talk HTTP, too.
- kaak3 6y agoUber recently blogged that they rebuilt the log analytics platform based on ClickHouse, replacing the previous ELK based one. The table schema choices made it easy to handle JSON formatted logs with changing schemas. https://eng.uber.com/logging/ https://eng.uber.com/logging/
- jetter 6y agoNice! Adding this to the post, thanks for the link!
- guardiangod 6y agoI am using Clickhouse at my workplace as a side project. I wrote a Rust app that dumps the daily traffic data collected from my company's products into a ClickHouse database. That's 1-5 billion rows, per day, with 60 days of data, onto a single i5 3500 desktop I have laying around. It returns a complex query in less than 5 minutes. I was gonna get a beef-ier server, but 5 minutes is fine for my task. I was flabbergasted.
- akudha 6y ago5 billion rows per day? What does your product do?
- guardiangod 6y agoSecurity products
- valiant-comma 6y agoRegarding #1 in the article, Elastic does have SQL query support[1]. I can’t speak to performance or other comparative metrics, but it’s worked well for my purposes. [1] https://www.elastic.co/guide/en/elasticsearch/reference/current/xpack-sql.html https://www.elastic.co/guide/en/elasticsearch/reference/curr...
- phillc73 6y ago> SQL is a perfect language for analytics. Slightly off topic, but I strongly agree with this statement and wonder why the languages used for a lot of data science work (R, Python) don't have such a strong focus on SQL. It might just be my brain, but SQL makes so much logical sense as a query language and, with small variances, is used to directly query so many databases. In R, why learn the data.tables (OK, speed) or dplyr paradigms, when SQL can be easily applied directly to dataframes? There are libraries to support this like sqldf[1], tidyquery[2] and duckdf[3] (author). And I'm sure the situation is similar in Python. This is not a post against great libraries like data.table and dplyr, which I do use from time to time. It's more of a question about why SQL is not more popular as the query language de jour for data science. [1] https://cran.r-project.org/web/packages/sqldf/index.html https://cran.r-project.org/web/packages/sqldf/index.html [2] https://github.com/ianmcook/tidyquery https://github.com/ianmcook/tidyquery [3] https://github.com/phillc73/duckdf https://github.com/phillc73/duckdf
- marcinzm 6y agoSQL tends to be non-composable which makes complicated scripts really messy to refactor and modify (or read even). CTEs make it more sensible but they're also a rather recent addition and don't fully solve the problem. Data Science tends to involve a lot of modifying of the same code rather than creating one off or static scripts.
- jfim 6y agoSQL doesn't compose all that well. For example, imagine that you have a complex query that handles a report. If someone says "hey we need the same report but with another filter on X," your options are to copy paste the SQL query with the change, create a view that can optionally have the filter (assuming the field that you'd want to filter on actually is still visible at the view level), or parse the SQL query into its tree form, mutate the tree, then turn it back into SQL. If you're using something like dplyr, then it's just an if statement when building your pipeline. Dbplyr also will generate SQL for you out of dplyr statements, it's pretty amazing IMHO.
- phillc73 6y ago
- proddata 6y agoIf you are looking an OSS ES replacement, CrateDB might also be worth a look :) Basically a best of both worlds combination of ES and PostgreSQL, perfect for time-series and log analytics.
- didip 6y agoDoes ClickHouse have integration with Superset and Grafana?
- daniel_levine 6y agoYes to both, Altinity maintains the ClickHouse Grafana plugin https://altinity.com/blog/2019/12/28/creating-beautiful-grafana-dashboards-on-clickhouse-a-tutorial https://altinity.com/blog/2019/12/28/creating-beautiful-graf... And Superset has a recommendation of a ClickHouse connector https://superset.apache.org/docs/databases/clickhouse https://superset.apache.org/docs/databases/clickhouse
- jcims 6y agoDoes ClickHouse or anything else out there that even remotely compete with Splunk for adhoc troubleshooting/forensics/threat hunting type work? I started off with Splunk and every time I try Elasticsearch I feel like I'm stuck in a cage. Probably why they can charge so much for it.
- supergirl 6y agowhy is splunk better than ES?
- jcims 6y agoI really want to answer you but I'm struggling a bit b/c I haven't worked with ES in a minute. Splunk just tends to be able to eat just about any kind of structured or unstructed content, operates on a pipeline concept similar to that of a unix shell or (gasp) powershell, and has a rich set of data manipulation of modification commands built in: https://docs.splunk.com/Documentation/SplunkLight/7.3.6/References/Searchcommandsbycategory https://docs.splunk.com/Documentation/SplunkLight/7.3.6/Refe... I primarily use it for security-related analysis, which is lowish on metrics and high on adhoc folding and mutilation of a very diverse set of data structures and types.
- eeZah7Ux 6y ago> ElasticSearch repo has jaw-dropping 1076 PRs merged for the same month Code change frequency is not a measure of quality or development speed. One organization can encourage bigger PRs while another encourage tiny, frequent changes. One can care about quality and stability while another can care very little about bugs.
- the-alchemist 6y agoAlso wanted to share my overall positive experience with Clickhouse. UPSIDES * started a 3-node cluster using the official Docker images super quickly * ingested billions of rows super fast * great compression (of course, depends on your data's characteristics) * features like https://clickhouse.tech/docs/en/engines/table-engines/mergetree-family/aggregatingmergetree/ https://clickhouse.tech/docs/en/engines/table-engines/merget... are amazing to see * ODBC support. I initially said "Who uses that??", but we used it to connect PostgreSQL and so we can keep the non-timeseries data in PostgreSQL but still access PostgreSQL tables in Clickhouse (!) * you can go the other way too: read Clickhouse from PostgreSQL (see https://github.com/Percona-Lab/clickhousedb_fdw https://github.com/Percona-Lab/clickhousedb_fdw, although we didn't try this) * PRs welcome, and quickly reviewed. (We improved the ODBC UUID support) * code quality is pretty high. DOWNSIDES * limited JOIN capabilities, which is expected from a timeseries-oriented database like Clickhouse. It's almost impossible to implement JOINs at this kind of scale. The philosophy is "If it won't be fast as scale, we don't support it" * not-quite-standard SQL syntax, but they've been improving it * limited DELETE support, which is also expected from this kind of database, but rarely used in the kinds of environments that CH usually runs in (how often do people delete data from ElasticSearch?) It's really an impressive piece of engineering. Hats off to the Yandex crew.
- hodgesrm 6y ago> It's really an impressive piece of engineering. Hats off to the Yandex crew. And thousands of contributors! Toward the end of 2020 over 680 unique users had submitted PRs and close to 2000 had opened issues. It's becoming a very large community.
- yamrzou 6y agoCould you share more details about the limited JOIN capabilities? AFAIK, Clickhouse has multiple join algorithms and supports on-disk joins to avoid out of memory: https://github.com/ClickHouse/ClickHouse/issues/10830 https://github.com/ClickHouse/ClickHouse/issues/10830 https://github.com/ClickHouse/ClickHouse/issues/9702#issuecomment-618299812 https://github.com/ClickHouse/ClickHouse/issues/9702#issueco...
- pachico 6y agoI've been using it successfully in production for year and a half. I can think of no other database that would give me real time aggregation over hundreds of millions of rows inserter every day for virtually zero cost. It's just a marvelous work.
- crb002 6y agoI wish they had a data store shoot-out like Techempower has for Web stacks.
- e12e 6y agoI just there was a foss loki-like solution built on ch - that was stable and used in production. I know there's a few projects (see below) - but I'm not aware of anything mature.. https://github.com/QXIP/cloki-go https://github.com/QXIP/cloki-go https://github.com/lmangani/cloki https://github.com/lmangani/cloki
- cduzz 6y agoAlmost nobody wants to use elasticsearch. People want to use kibana and put up with elasticsearch.
- dilyevsky 6y ago+1 we actually looked a ch as our debug logs backend and while it is great for the most part (it’s also incredibly memory hungry) kibana is really an es killer feature
- NicoJuicy 6y agoOr some people are considering it for eg. Facets
- mekster 6y ago> People want to use kibana and put up with elasticsearch. I don't buy this. It's just a mess of data dumps and it's not exactly providing focused experience. You need to take a full month to show what you want comfortably.
- wdb 6y agoI am curious how do they deal with GDPR or PPI when they do the logging? At first sight it looks like they are doing the logs themselves and not the API provider.