4 ms·
This is more a job ad than an article. I've wanted to try out Prometheus a while, but can't figure out how to: 1. make it highly available 2. play nice with f
by ymse 10y ago
This is more a job ad than an article. I've wanted to try out Prometheus a while, but can't figure out how to:
1. make it highly available
2. play nice with firewalls
If I deploy Prometheus outside a NAT, and want to monitor 100 physical machines on the inside with node_exporter, as well as a dozen different services, how to make these metrics available?
What if I have four identical NATed sites and want them all monitored by the same outside Prometheus instance(s)?
- XorNot 10y agoPrometheus can monitor via proxies. Or monitor via federated scraping. Though it is an odd requirement to need to only have Prometheus outside a firewall. Edit: guessing from your post - the way I'd do it is run a Prometheus instance at all 4 sites, and have them all federate each other. That way each site is the HA redundancy for the others.
- ymse 10y agoThat's sensible, thanks. What if the inside of the network is further segregated, and the Prometheus instance does not have access to all endpoints. Can I use the push gateway as a "proxy", or is a direct route required. Sorry for the stupid questions. Now it's getting entirely academic of course :) Edit: just re-read your post and saw the proxy note. Time to read the documentation again. This setup should make for a good blog post.
- bbrazil 10y agoIt's advised not to use the pushgateway in that fashion, it's for service-level batch jobs - not trying to subvert your organization's network security policies. See https://prometheus.io/docs/practices/pushing/ https://prometheus.io/docs/practices/pushing/
- bogomipz 10y agoCan you monitor alerts via a federated scraping? So for example you have one master alerts dashboard?
- XorNot 10y agoAlerts are push anyway at the moment - so sending multiple Prometheus alerts to a single alertmanager is easy. But yes, the actual "alert firing" metrics are exposed by the federated endpoints.
- bbrazil 10y agoYou'd lose annotations and other metadata doing that, plus it's another hop on the critical path of alerting. Probably not the wisest of ideas. You want your alerts coming from as close to what you're monitoring as possible.
- bbrazil 10y agoPrometheus developer here. To make it highly available, just run two identical Prometheus servers. That way if either fails the other is still sending alerts, and if both are working the Alertmanager will automatically deduplicate the alerts. As for networking, we recommend running Prometheus on the same network as what it's monitoring to avoid crossing failure domains. You could also open up whatever ports you need on the firewall, just as you would for any other application. Prometheus is no different from any other monitoring solution in requiring firewalls to allow it's traffic through.
- Millenis 10y agoSimilar experience here. We put a lot of time and effort into making our monitoring system more highly available than the thing it is monitoring. Not only that, only scaling vertically on a single node doesn't seem like a good design. When you're spread across many cloud providers, some on premise and a bunch of acquisition legacy stuff the polling and firewall opening becomes a blocker. There are ways to poll things and push metrics without opening millions of firewall ports to every security group. Sensu does that quite well, and it scales. I'm not pitching that as the saviour either as that has other trade-offs. However, Prometheus as far as I can see will suffer from the 2 points you make and all of the workarounds in terms of running duplicate identical nodes and federating them also come with various drawbacks. There are quite a few cloud providers who do an awesome job at providing those requirements. I think the only blocker on those is the current pricing model. As soon as somebody does something disruptive in this area I can see a great deal of convergence to it.
- bbrazil 10y ago> Not only that, only scaling vertically on a single node doesn't seem like a good design. For Prometheus at least, we're so efficient that it actually works out okay for the vast majority of users. You'd typically need thousands of instances doing the same thing inside a single datacenter before you get into our (admittedly more involved) horizontal sharding approach. http://www.robustperception.io/scaling-and-federating-prometheus/ http://www.robustperception.io/scaling-and-federating-promet... has more information. > There are ways to poll things and push metrics without opening millions of firewall ports to every security group. Sensu does that quite well, and it scales. I don't think that's quite a fair comparison. Sensu Just Works when there's no outbound firewall, Prometheus Just Works when there's no inbound firewall. If you add the other direction of firewall for either then things break down.
- Millenis 10y agoHi, honestly this feels like the conversation our team had a few years ago about Nagios. Some people were happy with pockets of monitoring servers dotted around with wiki pages full of links to get to them. It took a long time to clean that up and finally end up with a single central service. The outbound vs inbound firewalls is totally a fair comparison. At almost every company I have worked for the perimeter security blocks all incoming ports. Most places have no outbound port restrictions, or when they do they usually always have a proxy for traffic to go out (like https for updates etc). This is what makes the design of outbound only connections considerably better (Sensu, Datadog, insert pretty much any SaaS vendor). I honestly don't know of any company who would open up such an extreme number of ports required to allow scraping from an external monitoring tool. In a large enterprise you want to host the monitoring tool as a service for other groups which means potentially in a totally different data center or cloud and allowing a small list of subnets 'in' is considerably better than exposing every single server to external access. For the scaling I'm not sure I totally agree. There is a reason why distributed systems exist and that is to scale efficiently with some degree of redundancy baked in. Single node HA is probably the most inefficient method of scaling. It sounds like Prometheus needs to run properly on a single box for simplicity but over time needs to be broken up and made scalable beyond the bounds of a single server. Being limited to a single server in 2016, like Nagios was 10 years ago before they too started to split some things out, isn't something I'd advertise as a feature. I think right now Prometheus is probably filling the gap of what an industrial historian would do in a factory. It's a console you would have sitting next to the thing you were building that would provide real time measurements. We didn't want lots of individual consoles. We wanted a central large system that could hold years of historical data, as well as allowing people to query it in real time, and we liked the idea of not needing to double up on hardware costs. We're using a mix of elasticsearch and cassandra here to achieve that.