14 ms·
Grafana Mimir – Horizontally scalable long-term storage for Prometheus
- eatonphil 5y agoThere isn't a link to the project on the page (that I could find) so it almost looked like it's not open source. But here it is: https://github.com/grafana/mimir https://github.com/grafana/mimir.
- notamy 5y agoYou have to find the "Download" button and click it, it's very non-obvious :< The entire page seems to be designed to funnel you into signing up for their paid service, which makes sense, but still doesn't feel great...
- dewey 5y agoThe first CTA button on the page "Tutorial" links to a tutorial where the first step is to run the project with Docker. Doesn't really feel like an overly forced funnel to their paid service.
- candiddevmike 5y agoRecently switched from their cloud service back to on-premise. The cloud version wasn't being updated and the entire setup experience left a lot to be desired with how you connect their on-premise grafana agent, especially if you aren't using their easy button deployment stuff. Also, billing for metrics is insane, as on any given day my metric load may vary between 5-7k or more. This caused some operational overhead as I was constantly tweaking scrapers to reduce useless metrics. For $50/mo, you can self host everything easier, cheaper and with more control IMO.
- maccard 5y ago> For $50/mo, you can self host everything easier, cheaper and with more control IMO. Can you give an example as to how you could self host a grafana stack for $50/month? On AWS that buys you 4 cores, 8GB memory and 0 storage, and it's certainly not easier than clicking one button on the grafana website.
- krnlpnc 5y ago> $50/month? On AWS that buys you 4 cores, 8GB memory and 0 storage Self-hosting on AWS is kind of counterproductive. Look into "cloud" metal servers and the money will go much further.
- jrwr 5y agoTwo low end Hetzner/OVH Boxes for redundancy should do the trick
- jjeaff 5y agoThat's why AWS charges so much for outgoing traffic.
- alexjplant 5y agoThere are Helm charts available for all Grafana products so if you already run a Kubernetes cluster and have spare capacity you can just throw it up there. Loki supports shipping logs to GCS/S3 natively and Prometheus can use Cortex (also available as a Helm chart) to do the same. Once you throw Grafana behind SSO and implement a backup cronjob you're done until you reach scale and have to start deploying/scaling individual components separately. I implemented most of the above using Terraform on a managed DigitalOcean cluster on a Saturday a few months back; it wasn't super-hard. Alternatively you could rent a few VPSes someplace and use k3s or similar to get an unmanaged cluster.
- westurner 5y agoSuggestions for organizing a Helm + Terraform [+ k3s/k3d/MicroShift] provisioning and monitoring git repo with CI for job accounting? (without Ansible & AWX, which I'd create a role with for this too) - [ ] ENH,BLD: A cookiecutter for this would be cool
- m1keil 5y agoWe are running Grafana and Prometheus on a single t3.xlarge instance with 150GB gp3 EBS. Excluding traffic, it costs ~ $100 USD per month. We are doing 10 second scrapes and currently have roughly 141k active time series. In Grafana Cloud it would cost... 15000 metrics for free. 126000/1000 * $8 = $882 Now here's the real kicker.. the pricing Grafana puts on their website are assuming 60 second scrape interval (1 data point/minute or DPM). If you are doing 6 DPM, that's $8 * 6 per 1000 time series! So final bill.. drum rolls 126000/1000 * $8 * 6 = $6048 Yes. That's a 60x. Now, sure, we don't get the scale, the backups, the SLA.. but we can live without it. And when Prometheus will start acting slowly, we will just bump it to t3.2xl, or spend some time and filter out some of the noisy metrics we might have around. Btw, if you try to find any information about what is a "time series" or a "metric" on the Grafana's pricing page, good luck. https://grafana.com/docs/grafana-cloud/metrics-control-usage/control-prometheus-metrics-usage/changing-scrape-interval/ https://grafana.com/docs/grafana-cloud/metrics-control-usage...
- mdaniel 5y agoStill AGPL, which I guess makes sense given the rest of their stack is too: https://github.com/grafana/mimir/blob/mimir-2.0.0/LICENSE https://github.com/grafana/mimir/blob/mimir-2.0.0/LICENSE
- cett 5y agoPresumably AGPLv3 is why Grafana would rather develop this than Cortex?
- pracucci 5y agoHi. I'm Marco, I work at Grafana Labs and I'm a Grafana Mimir maintainer. We just published a couple of blog posts about the project, including more details on your question: https://grafana.com/blog/2022/03/30/announcing-grafana-mimir/ https://grafana.com/blog/2022/03/30/announcing-grafana-mimir... and https://grafana.com/blog/2022/03/30/qa-with-our-ceo-about-grafana-mimir/ https://grafana.com/blog/2022/03/30/qa-with-our-ceo-about-gr...
- cett 5y agoThank you for your answer. That seems like a reasonable strategy.
- halfmatthalfcat 5y agoHow does this stack up with https://github.com/thanos-io/thanos https://github.com/thanos-io/thanos, which I've used to pretty good success. The only criticism I have of Thanos though was the amount of moving pieces to maintain.
- pracucci 5y agoMimir has a microservices architecture. However, Mimir supports two deployment modes: monolithic and microservices. In monolithic mode you deploy Mimir as a single process and all microservices (Mimir components) run inside the same process. Then you scale it out running more replicas. Deployment modes are documented here: https://grafana.com/docs/mimir/latest/operators-guide/architecture/deployment-modes/ https://grafana.com/docs/mimir/latest/operators-guide/archit...
- netingle 5y ago(Tom here; I started the Cortex project on which Mimir is based and lead the team behind Mimir) Thanos is an awesome piece of software, and the Thanos team have done a great job building an vibrant community. I'm a big fan - so much so we used Thanos' storage in Cortex. Mimir builds on this and makes it even more scalable and performance (with a sharded compactor and query engine). Mimir is multitenant from day 1, whereas this is a relatively new thing in Thanos I believe. Mimir has a slightly different deployment model to Thanos, but honestly even this is converging. Generally: choosing Thanos is always going to be a good choice, but IMO choosing Mimir is an even better one :-p
- notacoward 5y agoMulti-tenancy is something that shouldn't be underestimated. A lot of people think it's just a checklist item until (a) they need it or (b) they try to implement it in an existing system. Kudos for making it a day-one feature.
- vladvasiliu 5y agoWhile I agree with your point in the general case, would you mind elaborating on the specific case of Prometheus? My understanding is that the recommended best-practice for Prometheus is to deploy as many of them as necessary, as close to the monitored infrastructure as possible. What use case would require deploying a single Mimir, so supposedly Prometheus (cluster) in the case of serving multiple tenants? Why not just deploy a dedicated Prometheus / Mimir stack per client?
- nosequel 5y agoGrafana Labs needs to make a convincing comparison chart of some kind between Mimir, Thanos, and Cortex. Thanos and Cortex are both mature projects and are both CNCF Incubating projects. Why would anyone switch to a new prometheus long-term storage solution from those? *EDIT*: I see from another reply there is a basic comparison to Cortex here: https://grafana.com/blog/2022/03/30/announcing-grafana-mimir/ https://grafana.com/blog/2022/03/30/announcing-grafana-mimir... To the Mimir folks, I'd love to see something similar Mimir v. Thanos.
- netingle 5y agoI agree! Which is why I put one in the blog post ;-) https://grafana.com/blog/2022/03/30/announcing-grafana-mimir/#introducing-grafana-mimir https://grafana.com/blog/2022/03/30/announcing-grafana-mimir...
- krnlpnc 5y agoI'm not seeing a comparison to Thanos
- alrlroipsp 5y agoWhy would you? Parent says its a comparison of Mimir and Cortex.
- krnlpnc 5y agoRe-read the full thread... >>Grafana Labs needs to make a convincing comparison chart of some kind between Mimir, Thanos, and Cortex. >I agree! Which is why I put one in the blog post ;-)
- smw 5y agoSeems like people should throw VictoriaMetrics into comparisons like this, as well?
- 5y ago
- eatonphil 5y agoIt's hard to tell exactly how this works but judging from the tutorial's docker-compose.yml [0] it looks like this runs as a separate API next to Prometheus and you tell Prometheus to write [1] to Mimir. I'm unclear how reads work from it or maybe there is no read? Maybe I'm completely misunderstanding. [0] https://github.com/grafana/mimir/blob/main/docs/sources/tutorials/play-with-grafana-mimir/docker-compose.yml https://github.com/grafana/mimir/blob/main/docs/sources/tuto... [1] https://github.com/grafana/mimir/blob/main/docs/sources/tutorials/play-with-grafana-mimir/config/prometheus.yaml#L23 https://github.com/grafana/mimir/blob/main/docs/sources/tuto...
- bboreham 5y agoIt’s a centralised multi-tenant store, supporting the Prometheus query API. So you can point clients directly at Mimir, they send in PromQL and they get data back in Json. (Note I work on Mimir)
- eatonphil 5y agoIs there an example of running mimir without prometheus?
- bboreham 5y agoFor example sending metrics from an OpenTelemetry pipeline. Mimir accepts the Prometheus remote-write api, which is protobuf-over-http; can be generated by anything really.
- k8sToGo 5y agoBut who does the scraping of the prometheus agents? Mimir or still prometheus server?
- bboreham 5y agoIf you have systems exporting metrics in Prometheus style, then you can use Prometheus to scrape them and remote-write to Mimir. You can alternately use Prometheus Agent, to save storing the data and running a query engine at the leaf. You can also use the OpenTelemetry suite to perform the same operation, though this is more appealing if you want some other OpenTelemetry features at the same time. Eg if you prefer the ‘pipeline’ style.
- ddon 5y agoLooks like an interesting alternative to Clickhouse with s3 backend...
- jhoechtl 5y agoWhat is the relationship to Loki?
- bboreham 5y agoSibling. Much of the architecture is similar; a number of components are shared in https://github.com/grafana/dskit https://github.com/grafana/dskit.
- dikei 5y agoSad news for Cortex, with most of the maintainer moving on to Mimir, I fear it's pretty much dead in the water.
- netingle 5y agoWe tried to address this question on the Q&A blog post: https://grafana.com/blog/2022/03/30/qa-with-our-ceo-about-grafana-mimir/#what-will-happen-to-the-cortex-project https://grafana.com/blog/2022/03/30/qa-with-our-ceo-about-gr... It doesn't have to mean the end for Cortex, but others will have to step up to lead the project. We've tried to put other maintainers in place to kick start this.
- sciurus 5y agoI was going to ask what the migration path was from Cortex to Mimir, but I see you've documented that at https://grafana.com/docs/mimir/latest/migration-guide/migrating-from-cortex/ https://grafana.com/docs/mimir/latest/migration-guide/migrat... . Thanks for the work you've done to make this easy.
- pracucci 5y agoThis video also shows a live migration from Cortex to Mimir (running in Kubernetes): https://www.youtube.com/watch?v=aaGxTcJmzBw&ab_channel=Grafana https://www.youtube.com/watch?v=aaGxTcJmzBw&ab_channel=Grafa...
- AndyNemmity 5y agoIf anything, this makes me less interested in moving from Thanos.
- cfors 5y agoMore engineering effort going into reinventing things that already exist to upsell people on Grafana cloud. What about focusing on the core value that Grafana provides, dashboards? Grafana 8 alerting is still in my opinion at a beta level. Dashboards as code has made no meaningful progress outside of community attempts in the past 3 years. The documentation for Grafana 8 alerts is still subpar. All of these things as a paid offering are more interesting than migrating my logging system or metrics system. Developers don't want to migrate their observability.
- detaro 5y agoIs there any competitor in the "primarily dashboards" space? Plenty things I know just use Grafana for small amounts of data where all this "5 new datastores!" isn't really useful, but dashboard improvements would be welcome.
- INTPenis 5y agoWhat issues have you seen with Grafana alerting? I'm curious because in my view it works so well that we abandoned alertmanager for Grafana alerts only well before v8.
- darkwater 5y agoHow did you define alarms as code in a practical way before v8? and after?
- INTPenis 5y agoDid not tbh. We have an ops department that do not complain about menial tasks. But of course IaC is the way we must follow.
- cfors 5y agoBuilding a dashboard by clickety/clacking around is not a menial task, consistency across dashboards is a a core unit of observability to ensure x-functional teams can discuss issues across a common language/viewpoint, which is only enforceable through a declarative dashboard syntax.
- MindTooth 5y agoHow does this compare to https://www.timescale.com/promscale https://www.timescale.com/promscale I’m looking into choosing a backend for my metrics and always open for suggestions.
- vineeth0297 5y agoHey! Promscale PM here :) Promscale is the open source observability backend for metrics and traces powered by SQL. Whereas Mimir/Cortex is designed only for metrics. Key differences: 1. Promscale is light in architecture as all you need is Promscale connector + TimescaleDB to store and analyse metrics, traces where as Cortex comes with highly scalable micro-services architecture this requires deploying 10's of services like ingestor, distributor, querier, etc. 2. Promscale offers storage for metrics, traces and logs (in future). One system for all observability data. whereas the Mimir/Cortex is purpose built for metrics. 3. Promscale supports querying the metrics using PromQL, SQL and traces using Jaeger query and SQL. whereas in Cortex/Mimir all you can use is PromQL for metrics querying. 4. The Observability data in Cortex/Mimir is stored in object store like S3, GCS whereas in Promscale the data is stored in relational database i.e. TimescaleDB. This means that Promscale can support more complex analytics via SQL but Cortex is better for horizontal scalability at really large scales. 5. Promscale offers per metric retention, whereas Cortex/Mimir offers a global retention policy across the metrics. I hope this answers your question!
- tarun_anand 5y agoThanks... how do we do reporting/dashboards/alerts with Promscale? Also, any performance benchmarks?
- vineeth0297 5y agoPromscale supports reporting/ingestion of data using Prometheus remote-write for metrics, OTLP (OpenTelemetry Line Protocol) for traces. Dashboards you can use Promscale as Prometheus datasource for PromQL based querying, visualising, as Jaeger datasource for querying, visualising traces and as PostgreSQL datasource to query both metrics and traces using SQL. If you are interested in visualising data using SQL, we recently published a blog on visualising traces using SQL (https://www.timescale.com/blog/learn-opentelemetry-tracing-with-this-lightweight-microservices-demo/ https://www.timescale.com/blog/learn-opentelemetry-tracing-w...) Alerts needs to be configured on the Prometheus end, Promscale doesn't support alerting at the moment. But expect the native alerting from Promscale in the upcoming releases. We have internally tested Promscale at 1Mil samples/sec, here is the resource recommendation guide for Promscale https://docs.timescale.com/promscale/latest/installation/recomm-guide/ https://docs.timescale.com/promscale/latest/installation/rec... If you are interested in evaluating, setting up Promscale reach out to us in Timescale community slack(http://slack.timescale.com/ http://slack.timescale.com/) in #promscale channel.
- Thaxll 5y agoSo many solutions to the same problem, how does it compare to Victoria Metrics?
- outsb 5y agoGiven Victoria Metrics is the only solution I've seen to make data comparing it to other systems easily accessible as part of official documentation, it's the only one I pay attention to. I knew from reading the docs what VM excelled at and areas it was weak in, long before I ever ran it (and expectations from running it matched the documentation). I hate aspirational marketing-saturated campaigns for deep tech projects where standards should obviously be higher, it speaks more about intended audience than it does the solution, and that's why in this respect VM is automatically a cut above the rest.
- cip01 5y agoCortex, Thanos and Mimir all support "remote-read" protocol (documented in Prometheus: https://prometheus.io/docs/prometheus/latest/storage/#remote-storage-integrations https://prometheus.io/docs/prometheus/latest/storage/#remote...), so external systems (eg Prometheus) can read data from them easily.
- valyala 5y agoIt would be great if you could provide a few practical examples for "Prometheus remote-read" protocol given its' restrictions [1]. [1] https://github.com/prometheus/prometheus/issues/4456 https://github.com/prometheus/prometheus/issues/4456
- cip01 5y agoWhich restrictions do you have in mind? Quick look at the issue looks like it wanted to avoid using local storage by Prometheus, but that’s Prometheus specific problem, not remote-read problem. Remote-read is a generic protocol (https://github.com/prometheus/prometheus/blob/a1121efc18ba15e3dce48042639648afb83114e2/prompb/remote.proto#L31 https://github.com/prometheus/prometheus/blob/a1121efc18ba15...), you pass query (start/end time and matchers), and get back data.
- SuperQue 5y agoOne interesting question I have is regards to global availability. With our current Thanos deployment, we can tie a single geo regional deployment together with a tiered query engine. Basically like this: "Global Query Layer" -> "Zone Cluster Query Layer" -> "Prom Sidecar / Thanos Store" We can duplicate the "Global Query Layer" in multiple geo regions with their own replicated Grafana instances. If a single region/zone has trouble we can still access metrics in other regions/zones. This avoids Thanos having any SPoFs for large multi-user(Dev/SRE) orgs.
- bboreham 5y agoThe typical way to run Mimir is centralised, with different regions/datacenters feeding metrics in to one place. You can run that central system across multiple AZs. If you run Mimir with an object store (e.g. S3) that supports replication then you can have copies in multiple geographies and query them, but the copies will not have the most recent data. (Note I work on Mimir)
- ddreier 5y agoThis is one of my favorite things about Thanos. We run Prometheus in multiple private datacenters, multiple AWS regions across multiple AWS accounts, and multiple Azure regions across multiple subscriptions. We have three global labels: cloud, region, and environment. With Thanos's Store/Querier architecture we have a single Datasource in Grafana where we can quickly query any metric from any environment across the breadth of our infrastructure. It's really a shame that Loki in particular doesn't share this kind of architecture. Seems like Mimir, frustratingly, will share this deficiency.
- misiti3780 5y agoWhat is the best SASS based dashboard solution for Prometheus?
- heinrichhartman 5y agoGrafana Cloud
- misiti3780 5y agothanks
- deleted 5y ago[deleted]
- sriv1211 5y agoWhat's the latency between sending a metric and being able to query it when using object storage (s3) instead of block storage? How do the transfer/retrieval (GET/PUT) costs factor in as well?
- pracucci 5y agoGood question! Grafana Mimir guarantees read-after-write. If a write request succeed, the metric samples you've written are guaranteed to be queried by any subsequent query. Mimir employes write deamplification: it doesn't write immediately to the object storage but keeps most recently written data in-memory and/or local disk. Mimir also employes several shared caches (supports Memcached) to reduce object storage (S3) access as much as possible. You can learn more here in the Mimir architecture documentation: https://grafana.com/docs/mimir/latest/operators-guide/architecture/about-grafana-mimir-architecture/ https://grafana.com/docs/mimir/latest/operators-guide/archit...
- bbu 5y agoi don't get why there's so much hate here. cortex is a pain to configure and maintain. would be awesome to have mimir address these issue!
- firstSpeaker 5y agoHow does it work with Rules? So far I cannot see if this can be a replacement for prometheus since I cannot see how can we re-use our prometheus rules with Mimir. Anyone knows anything around that?
- pracucci 5y agoMimir includes a ruler component, which is responsibile to evaluate Prometheus recording and alerting rules. It also exposes a set of APIs to configure the rule groups. For example, you can use this API to upload a rule group: https://grafana.com/docs/mimir/latest/operators-guide/reference-http-api/#set-rule-group https://grafana.com/docs/mimir/latest/operators-guide/refere... Mimir is released with a CLI tool called "mimirtool" which, among other things, allow you to configure the rule groups (under the hood, it calls the Mimir API). Mimirtool documentation is here: https://grafana.com/docs/mimir/latest/operators-guide/tools/mimirtool/ https://grafana.com/docs/mimir/latest/operators-guide/tools/...
- firstSpeaker 5y agoThank you for the reply.
- monstrado 5y agoIs this the project you guys referenced using Apache Arrow for?
- netingle 5y agoI don't think so! I think thats being used in Tempo, but I'm not sure.
- number101010 5y agoWe are definitely investigating columnar formats in Tempo to store traces. We expect it to drastically accelerate search as well as open up more complex querying and eventually metrics from distributed tracing data. However, we are currently primarily targeting Parquet as our columnar format in object storage. Expect an announcement soon!
- bboreham 5y agoMaybe you're thinking of this - the data structure used by datasources for Grafana dashboards: https://grafana.com/docs/grafana/latest/developers/plugins/data-frames/#apache-arrow https://grafana.com/docs/grafana/latest/developers/plugins/d...
- mgarciaisaia 5y agoThe thing I need most right now is a confirmation that it's named after this tweet: https://twitter.com/mmoriqomm/status/1272552214658117638 https://twitter.com/mmoriqomm/status/1272552214658117638
- young_unixer 5y agoCoincidentally, "mimir" is a funny, baby-like way of saying "dormir" (to sleep) in Spanish.
- vladsanchez 5y agoSo true!!! LOL I related to "Vamos a mimir!" when I read it!!! ROFL
- estebarb 5y agoTechnical meetings are going to be fun with hispanic devs... "And finally we sent the metrics to Mimir /giggles/" Sadly they don't support encryption at rest (sorry, I really had to do one more pun)
- camel_gopher 5y ago"the most scalable open source TSDB in the world" You can be scalable, and still cost a lot of money to scale out. Unit economics are important.
- mr-karan 5y agoFolks looking for a solution to storing Prometheus metrics from multiple places, definitely consider exploring Victoriametrics. I'm running a single Victoriametrics instance which has 230bn metrics, consuming ~4GB of memory and barely 200m of CPU utilization (only spikes to ~1.5cores when it flushes these datapoints from RAM to disk). I've previously[1] shared my experience of setting up Victoriametrics for long term Prometheus storage back in 2020 and since then this product has just kept getting better. Over time, I switched to `vmagent` and `vmalert` as well which offer some nice little things (like did you know, you can't break up the scrape config of Prometheus into multiple files? `vmagent` does that happily). The whole setup is very easy to manage for an Ops person (as compared to Thanos/Cortex. Yet to checkout Mimir though!) as well. I've barely had to tweak any default configs that come in Victoriametrics and I even increased the retention of metrics from a month to multiple months after gaining confidence in prod. [1]: https://zerodha.tech/blog/infra-monitoring-at-zerodha/ https://zerodha.tech/blog/infra-monitoring-at-zerodha/
- jaigupta 5y agoThis is about Prometheus but Mimir makes it interesting. I can't find any other open source time series database except Mimir/Cortex which allows this much scale (clustering options in their open source version). Our use case will have high cardinality and Mimir seems to fit very well. Can we use Prometheus/Mimir as general purpose time series database? Prometheus is built for monitoring purposes and may not be for general purpose time series databases like InfluxDB (I am hoping to be wrong). What are the disadvantages/limitations for using Prometheus/Mimir as general purpose time series database?
- valyala 5y ago> I can't find any other open source time series database except Mimir/Cortex which allows this much scale (clustering options in their open source version) The following open source time series databases also can scale horizontally to many nodes: - Thanos - https://github.com/thanos-io/thanos/ https://github.com/thanos-io/thanos/ - M3 - https://github.com/m3db/m3 https://github.com/m3db/m3 - Cluster version of VictoriaMetrics - https://docs.victoriametrics.com/Cluster-VictoriaMetrics.html https://docs.victoriametrics.com/Cluster-VictoriaMetrics.htm... (I'm CTO at VictoriaMetrics) > Can we use Prometheus/Mimir as general purpose time series database? This depends on what do you mean under "general purpose time series database". Prometheus/Mimir are optimized for storing (timestamp, value) series where timestamp is a unix timestamp in milliseconds and value is a floating-point number. Each series has a name and can have arbitrary set of additional (label=value) labels. Prometheus/Mimir aren't optimized for storing and processing series of other value types such as strings (aka logs) and complex datastructures (aka events and traces). So, if you need storing time series with floating-point values, then Prometheus/Mimir may be a good fit. Otherwise take a look at ClickHouse [1] - it can efficiently store and process time series with values of arbitrary types. [1] https://clickhouse.com/ https://clickhouse.com/
- jaigupta 5y agoI meant all Prometheus based solutions, includes Thanos, M3, VictoriaMetrics. Thank you for your answer.
- ankitnayan 5y agoI would love to see some benchmarks when making such a heavy claim. I would be interested in knowing performance of ingestion rate, query timings and resource usage.