3 ms·
> Umm, yeah, but well, isn't that the punchline of every monitoring project? ;-) No actually :) From reading countless websites for the multitude of monitorin
by bbrazil 10y ago
> Umm, yeah, but well, isn't that the punchline of every monitoring project? ;-)
No actually :)
From reading countless websites for the multitude of monitoring solutions out there, the standard pitch is that they'll make your operations more efficient, magically detect problems and give you new insights into your systems.
It's actually really annoying, as you have to dig deep into their docs and squint at screenshots to figure out that actually it's just another Nagios clone :)
> I don't understand. Isn't that twitter blog essentially saying that they went from pull to push? Or is the main point here that they changed to have the agent on each host that collects everything on that host before transferring those metrics to the server (whether that's push or pull)?
What I take from it was had they kept the same design and just changed to push, they'd have continued to have had the exact same problems. I'd view the addition of isolation and health information as what really solved the issue, and whether push or pull is better for that is a wash.
> But I'm not sure Prometheus bends to this either, since IIUC each metric is stored in a separate file, and if we'd have a bunch of metrics for each job ID in the system, this wouldn't really work out, would it?
I'd need a bit more detail to give an exact solution, but Prometheus should work fine for for this.
Metrics are fetched over HTTP by Prometheus, and they usually don't come from files. You'd either have Prometheus scrape each job individually, or scrape your controller daemon that'd attach a labels indicating the job of each metric.
- jabl 10y ago> Metrics are fetched over HTTP by Prometheus, and they usually don't come from files. You'd either have Prometheus scrape each job individually, or scrape your controller daemon that'd attach a labels indicating the job of each metric. I was thinking more of the data model (https://prometheus.io/docs/concepts/data_model/ https://prometheus.io/docs/concepts/data_model/). Specifically, from https://prometheus.io/docs/practices/naming/ https://prometheus.io/docs/practices/naming/ , "Remember that every unique key-value label pair represents a new time series, which can dramatically increase the amount of data stored. Do not use labels to store dimensions with high cardinality (many different label values), such as user IDs, email addresses, or other unbounded sets of values.". So in our case the job ID would be such an unbounded value..?
- bbrazil 10y agoThese would be the exact details I'd need to advise you. It depends on how many jobs you have, how much churn there is, and what exactly you want to monitor. If jobs are short-lived then tracking individual jobs would be unwise, something like the ELK stack intended for event logging would be better. If jobs are long-lived and there's not many of them then you should be okay. Otherwise you'd just be looking at tracking system rather than per-job stats. To give a very rough idea, if you can keep it below say 10M metrics across the history a single Prometheus server has that should be okay with the current implementation.