4 ms·
I often encounter a lot of confusion about push vs. pull and whether you should pick Prometheus or Influx. Prometheus comes along every so often and scrapes me
by 101km 9y ago
I often encounter a lot of confusion about push vs. pull and whether you should pick Prometheus or Influx.
Prometheus comes along every so often and scrapes metrics your program exposes via a very simple API. This has the advantage that your code doesn't need to know about some endpoint of some cluster, it just needs to buffer up and expose some info it knows about itself from the recent past.
You don't need to maintain a cluster of Prometheus either, you can just run more than one for redundancy (kind of like an active-active) - it is meant for relatively ephemeral information, it is efficient, and one big Prometheus node will probably do you fine.
Where to store compacted historical metrics (less coarse resolution, still interesting data that you might not want to throw away) and how has been an open question. Sinking it into influx could be a good answer, so this is welcome news.
- pauldix 9y agoThat's exactly what we're going for. Ideally, we'll also expose PromQL query functionality in Influx so Prometheus users can hit their historical data in Influx directly just like they would with Prometheus. Still early days on all this work, but I was very inspired coming away from PromCon last month.
- pstuart 9y agoConversely, one has to tell the collector about a new device to poll. Being able to fire off a UDP packet with the metric and move on seems like would be the most lightweight approach.
- sytse 9y agoTwo things to consider: 1. UDP packets are lossy, we had the UDP buffers of our Influx server fill up, and it took us a long time to detect we were dropping packets. 2. Many people want to detect when they are not getting data from an endpoint. Polling is a great way to quickly detect a endpoint is down.
- pstuart 9y agoFor argument's sake (abuse is down the hall): 1. My understanding of UDP being lossy typically refers to it happening via transit but your example is an endpoint failure. 2. Since the whole point of metrics is to keep track of operations then the monitoring of the metrics themselves should be alerting to anomalies?
- 101km 9y agoIt sure is until you consider that now you have to maintain that endpoint and have it be constantly up or it will simply miss metrics. Not to mention the occasional network partition is outside your control. Additionally, I would argue some sort of discovery/registry mechanism is in order anyway. For example, Prometheus has very solid integration with Kubernetes. In this universe you have one central control plane thingy (the scheduler) responsible for bringing resources up and down, it updates the collector (Prometheus) about all the devices (pods, containers, damn we have too many terms for things) accordingly, which in turn scrapes on a best effort basis. Once everything is hooked up like this you're basically guaranteed that applications that provide information about themselves will get scraped eventually. If they are reachable you'll know what is going on inside them, and if they are not, you'll know that too. The advantages of this sort of decoupling are subtle and difficult to get right, but you'll be thankful down the line for having done so correctly from the get go.
- bbrazil 9y agoThere's two issues with that, first the statsd event approach has scaling problems as relative to an instrumentation event a UDP packet is actually quite heavyweight (https://www.robustperception.io/which-kind-of-push-events-or-metrics/ https://www.robustperception.io/which-kind-of-push-events-or...) and secondly how do you know that all devices that are meant to be sending are sending (https://www.robustperception.io/push-needs-service-discovery/ https://www.robustperception.io/push-needs-service-discovery...).
- fh973 9y agoWhen you need to consider scalability, pull also degrades more nicely.