8 ms·
Amazon Managed Service for Prometheus
- pickledish 6y ago14 cents per "query processing minute" sounds like it could add up very fast. Prom queries can get somewhat complex and it's not rare at all IME to have a dashboard making several multi-second queries per load (whether that falls into "you're using Prometheus wrong" being a separate discussion of course) Edit: The example from their pricing page: > We will assume you have 1 end user monitoring a dashboard for an average of 2 hours per day refreshing it every 60 seconds with 20 chart widgets per dashboard (assuming 1 PromQL query per widget)... assuming 18ms per query for this example. Comes out to over $3 per month in query costs. Replace this 1 person with a TV showing the dashboard all day, and the cost jumps to $36, for just one dashboard and (again IME) overly fast query estimates... o.O
- gravypod 6y agoDoes it put any limits on cardinality of metrics? Grafana cloud's offering was absolutely awful for my use cases. They charge per-series so if you have metrics with a "pod=..." label your prices go through the roof.
- heliodor 6y agoPlenty has been written about not using the server/container/pod id as a label because it leads to high cardinality which leads to poor performance (cost aside). Time series databases have been purpose-built for certain workloads and you can consider this their weakness.
- gravypod 6y agoPlenty has also been written about the bugs/issues that have cropped up that are only visible when inspecting what regions/nodes/cgroups an issue is coming from [0]. My use case wasn't exactly `pod=...` but it was very similar. It was more like `device=...`. Also, for a huge application, it's not uncommon to have 100s or even 1000s of metrics that are important to application health/performance. Constantly saying "do you really need X? It will cost us Y" will lead to an extremely under-monitored application. [0] - https://cloud.google.com/blog/products/management-tools/sre-keeps-digging-to-prevent-problems https://cloud.google.com/blog/products/management-tools/sre-...
- heliodor 6y agoPlenty of companies run their own servers because cloud is too expensive at their scale. Same goes for metrics. It's a direct result of one-price-fits-all pricing models for software as well as pricing that is not correctly tied to value.
- webo 6y agoI like Weave Cloud’s Prometheus hosting model — it’s per host, which is predictable and forecastable.
- kasey_junk 6y agoEvery managed metrics system will put a limit on cardinality because all mainstream available metrics systems cost more per cardinality to query and store. If they don’t limit that you can assume you or some other customer is going to use up the clusters resources and cause an outage. Like most metrics systems, under the covers in Prometheus each unique combination of dimensions is the same as a new metric line.
- deleted 6y ago[deleted]
- valyala 6y agoGrafana cloud sets high prices for high-cardinality metrics because the underlying system - Cortex - isn't optimized well for storing high number of unique time series. For example, it requires at least 15GB of RAM for processing a million of active time series [1]. This means high infrastructure costs, which increase pricing for end users. Other systems such as VictoriaMetrics require up to 15x lower RAM for the same metrics' cardinality [2]. [1] https://github.com/cortexproject/cortex/blob/67648aabae70f1948b93b90ee56a96c1d7558422/docs/guides/capacity-planning.md https://github.com/cortexproject/cortex/blob/67648aabae70f19... [2] https://victoriametrics.github.io/#capacity-planning https://victoriametrics.github.io/#capacity-planning
- edoceo 6y agoNow do six dashboards, 10 widgets each, multiple viewers, 18h/day and one slowish query on each dashboard. Seems like we get to hundred+ pretty quick
- bboreham 6y agoCaching means that multiple viewers cost very little extra. (I am a Cortex maintainer)
- pram 6y agoYeah I dunno about this, and the grafana service. They’re not exactly complicated to run on their own. At this pricing you may as well be on Datadog.
- stevekemp 6y agoThis seems more interesting of the two, grafana is pretty simple to setup and maintain. The harder part is handling the metrics themselves, be it with influxdb, prometheus, or something else.
- markcartertm 6y agosetting up one Prometheus server is easy. scaling, HA, Metrics retention for more than 3 days not so much.
- heliodor 6y agoLook at VictoriaMetrics (and the related products vmalert and vmagent) for a much easier and pleasant experience as a drop-in Prometheus replacement.
- nrmitchi 6y agoI've commented fairly heavily in the related Grafana thread. Prometheus is a bit of a different story. It does have some operational overhead when you get to a certain point, and scaling it out is not always trivial. Assuming it works, there is value-add on this one, and the pricing is more in line with active use (ie, a cost+ model, which is more typical of AWS services)
- deleted 6y ago[deleted]
- valyala 6y agoAmazon Managed Service for Prometheus is based on Cortex. It is quite expensive in terms of operational and infrastructure costs compared to VictoriaMetrics [1] according to case studies from VictoriaMetrics users [2]. This may explain quite high costs for AMP. [1] https://victoriametrics.github.io/FAQ.html#what-is-the-difference-between-victoriametrics-and-cortex https://victoriametrics.github.io/FAQ.html#what-is-the-diffe... [2] https://victoriametrics.github.io/CaseStudies.html https://victoriametrics.github.io/CaseStudies.html Disclaimer: I'm core developer of VictoriaMetrics, so feel free asking any questions about it or about our competitors :)
- WoahNoun 6y agoEveryone here complaining about the pricing on the managed Grafana and Prometheus services have clearly never worked at a shop using SumoLogic. Log/metric processing/querying is expensive for a reason.
- eminence32 6y agoFrom the pricing section: > AMP counts each metric sample ingested to the secured Prometheus-compatible endpoint. AMP also calculates the stored metric samples and metric metadata in gigabytes (GB), where 1GB is 230 bytes. Surely that's a typo, right?
- biot 6y agoLikely a casualty of copy and paste that left out the superscript formatting. 1GB is 2^30 bytes.
- vishuk 6y agoDo we know which scalable prometheus backend are they running? Chronosphere? Thanos?
- bmurphy1976 6y agoThe Grafana blog post mentions Cortex, something I'm not familiar with: https://grafana.com/blog/2020/12/15/announcing-amazon-managed-service-for-grafana/ https://grafana.com/blog/2020/12/15/announcing-amazon-manage...
- bboreham 6y agoIt’s Cortex, though the particular configuration shares a lot of code with Thanos. (I am a Cortex maintainer)
- hagen1778 6y agoIf you know technical details, are there any metrics cardinality limitations?
- bboreham 6y agoThere are soft limits _everywhere_, to stop people shooting themselves in the foot. Those can be raised by admins after checking the user knows what they are doing. I do not know what the practical limits are right now; especially I do not know what size hardware AWS run it on. If you were to search the Cortex Slack you would find people talking about instances with 100 million series, also people talking about work to improve scalability.
- valyala 6y agoThey run Cortex. Cortex architecture looks a bit over-engineered [1] when compared to other backends for Prometheus [2]. This may negatively affect system reliability while increasing operational costs and infrastructure costs. I hope Cortex architecture will be simplified in the future. [1] https://github.com/cortexproject/cortex/blob/master/docs/architecture.md https://github.com/cortexproject/cortex/blob/master/docs/arc... [2] https://victoriametrics.github.io/Cluster-VictoriaMetrics.html#architecture-overview https://victoriametrics.github.io/Cluster-VictoriaMetrics.ht...
- latchkey 6y agoI just went through the "process" of installing Grafana, Loki, Promtail and Prometheus on an ubuntu box and it is almost like the company behind all of this has gone out of the their way to make it hard. It isn't really _that_ difficult to get set up, but it also isn't 'apt install' easy (you really want me to create my own startup scripts?) and required me to build my own documentation on how I installed everything.
- john_moscow 6y agoIt's almost like the company behind it wants to see some profit after pouring millions of dollars into developing these tools. Except, in 2020 you cannot just have a closed-source easy-to-use documented and supported product with a license fee. Not in the server market, at least. Everything must be free and open-source, and you are expected to make money by offering a hosted service. Except, good luck competing with Big Cloud.
- RocketSyntax 6y agoIt's extremely worrisome. The incentive to spend your early mornings, nights, and weekends building something awesome to free yourself from corporate life is fading away. They need to institute some kind of royalty program or at least dedicate engineers to helping maintain the projects they make into services. Almost have to change gears and get into a scientific field that isn't computer science.
- rfratto 6y agoOne of the Loki maintainers here (though I mostly work on other stuff now). I promise it's not difficult on purpose. We've put a lot of effort into optimizing the Kubernetes experience that non-containerized installations haven't been getting as much attention. We'd be thrilled to have system packages for Loki that also set it up as a service, it's just not something we've been able to spend time doing ourselves yet.
- latchkey 6y agoIt isn't just loki, but the whole stack. Grafana is the only project mentioned that has a debian installer. The expectation that someone doing greenfield development is going to jump into k8s just to use the software is kind of weird.
- alexhf 6y agoI don't see any mention of Pushgateway. They'll need to add that or I won't be able to monitor ephemeral jobs.
- mchene 6y agoHey... Marc here from AWS. I'm the PM lead for this service. Thank you for the feedback. Pushgateway is important for our customers and it is a feature we are looking to support as part of our roadmap. For the time being, you can continue to use the Pushgateway as you do today and remote write the metrics to AMP for long term storage and querying!
- slyall 6y agoThe pricing just for the ingest seems way off. $0.002 for 10,000 metrics might not seem like much by even a simple node_exporter will grab 700 metrics every 15 seconds. Thats $24/month just to ingest the cpu/ram/diskspace data from each server. Plus storage and query costs. At work I have a single r4.xlarge instance handling 1.3 million metrics every 15 seconds. Storage is not clustered but cost is only $500/month. It would cost me $45k/month just for the ingest with the new managed service.
- mchusma 6y agoTheir pricing for these managed services used to be "no brainer" (something like the cost of compute only, or maybe a <30% upcharge). Managed airflow was similarly very expensive (maybe 3x the cost). Just not worth it. Bummer.
- wpietri 6y agoYeah, it turns out there's a lot of money to be made from people who don't have a good grasp of the fundamentals. We got a marketing email from Huggingface recently about their ML-models-as-a-service offering: https://huggingface.co/pricing https://huggingface.co/pricing One of my colleagues asked if it might be better than creating our own infrastructure for that. I ran the numbers for one of our recent jobs, feeding a million tweets to two ML models to see which worked better. That would have cost about $1800 on Huggingface. Using AWS spot instances, it was maybe $25 for us to run ourselves. Of course, we can do it at that price because we are paying for engineers and plan on classifying enormous amounts of text, so it works out for us. Plenty of other people probably should just use Huggingface. But I can't help looking at that 70x markup and think, "Fuck me? No, fuck you!"
- throwaway343233 6y agoPricing makes sense if you consider how Amazon operates at this point. You put basically a MVP product out there with abnormal pricing. Your enterprise customers that are drowning in money can start using it and using that money you can grow your org by hiring more engineers. At this point you start working on adding new features and do cost optimization. Since your whole architecture was designed based on "we have to ship this ASAP", you deliver some real nice cost reduction easily. Then you reflect this to your customers and gain goodwill and good PR.
- hagen1778 6y agoI wonder if it will be possible to migrate your data somewhere else once it becomes too expensive.
- backing 6y agoHoe can I hide all Amazon and Google news on HN ? Do you know an alternative of HN without big tech lobby? Thanks.
- joana035 6y agoI'm interested in this too, more I see aws dominating every aspect of our life, more depressed I become.
- aluminussoma 6y agoI very much dislike Prometheus, but the fact that AWS is offering it as a managed service means I am in the minority. I attribute much of Prometheus' success to the influence of ex-Googlers. They joined other companies, had a lot of clout, and sought out a tool that was similar to what they once used. I understand that the Google version of Prometheus is deprecated but there is no commercial equivalent.
- dilyevsky 6y agoWhat is in your opinion a better open source alternative to prometheus? Borgmon was inspiration for prometheus but was a totally different project so it is a complete rewrite
- codeduck 6y agoI feel like a broken record, but we are having great success with Victoria metrics as a drop in replacement.
- heliodor 6y agoPromscale looks interesting. Keep the architecture of Prometheus while storing the data in TimescaleDB and using SQL as the query language (together with the TimescaleDB-specific extensions to it). Does anyone actually like PromQL?
- notesinafield 6y agoIts nearly weekly we bump up against the limits of timeseries aggregation. Id take anything else foss at this point.
- AYBABTME 6y agoI wonder how AWS is supporting the development of Prometheus. Are they financing the OSS developers who are spending countless hours dedicated to the project?
- bboreham 6y agoAWS is an investor in Weaveworks where the implementation (Cortex) was first created. Weaveworks had two Prometheus maintainers on staff at the time. In the announcement it says AWS have a commercial relationship with Grafana Labs, where several Prometheus maintainers, community managers, etc. currently work. (I work for Weaveworks)
- AzzieElbab 6y agoNow that Aws ate the world, can we get some useable gui or consistent cli?