4 ms·
> VictoriaMetrics :( Not actually Prometheus-compatible, sloppy code, spotty docs. I have no idea why this dumb product continues to attract users. https://p
by sagichmal 5y ago
> VictoriaMetrics
:(
Not actually Prometheus-compatible, sloppy code, spotty docs. I have no idea why this dumb product continues to attract users.
https://prometheus.io/blog/2021/05/04/prometheus-conformance-remote-write-compliance/ https://prometheus.io/blog/2021/05/04/prometheus-conformance...
> Telegraf, which is to metrics sort of what Logstash is to logs: a swiss-army knife tool that adapts arbitrary inputs to arbitrary output formats. We run Telegraf agents on our nodes to scrape local Prometheus sources, and Vicky scrapes Telegraf. Telegraf simplifies the networking for our metrics; it means Vicky (and our iptables rules) only need to know about one Prometheus endpoint per node.
Normally you just use a regular Prometheus server to do this. Why add another, different technology to the stack?
> We spent some time scaling it with Thanos, and Thanos was a lot, as far as ops hassle goes.
It really isn't -- assuming you're not trying to bend Prometheus into something it isn't. Prometheus works using a federated, pull-based architecture. It expects to be near the things it's monitoring, and expects you to build out a hierarchy of infrastructure, in layers, to handle larger scopes.
This is structurally different to what I'll call the "clustering" model of scale, where you have all your data sources pushing their data, aggregating maybe on the machine or datacenter level, but then shuttling everything to a single central place, which you scale vertically from the perspective of your users. This appears to be what you want to do, based on the prevalence of push-based tech in your stack.
Prometheus doesn't work this way. Some people really want it to work this way, and have even created entire product lines that make it look as if it works this way (Cortex, M3db) but it's fundamentally just not how it's designed to be used. If you try to make it work this way yourself, you'll certainly get frustrated.
- mrkurt 5y ago> Normally you just use a regular Prometheus server to do this. Why add another, different technology to the stack? Our physical hosts have hundreds of services exporting metrics. And many of those exported metrics are from untrusted sources. So we can both rewrite labels and decrease the scrape endpoint discoverability problem by aggregating them in one place. > Not actually Prometheus-compatible, sloppy code, spotty docs. I have no idea why this dumb product continues to attract users. Because it works incredibly well, it's easy to operate, and handles multi tenancy for us.
- sagichmal 5y ago> Our physical hosts have hundreds of services exporting metrics. And many of those exported metrics are from untrusted sources. So we can both rewrite labels and decrease the scrape endpoint discoverability problem by aggregating them in one place. OK, but Prometheus can do all of this just fine? > Because it works incredibly well, it's easy to operate, and handles multi tenancy for us. Again, Prometheus itself ticks all of these boxes, too, if you're not trying to force it to be something it's not.
- tptacek 5y agoWe're not forcing Prometheus to be anything, since we're not using it. What Prometheus wants to be is not really a relevant constraint in our design space. A topologically simple, scalable, multi-tenant cluster that presents as just a giant bucket of metrics to our users is what we wanted, and we got it. There's an interesting discussion to be had about how our infrastructure works; for example, in the abstract, I'd prefer a "pure" pull-based design too. But things appear and disappear on our network a lot, and remote write simplifies a lot of configuration for us, so I don't think it's going anywhere. I think you're reading a critique of Prometheus that isn't really present in what we're writing. Prometheus is great! Everyone should use it! Our needs are weird, since we're handling metrics as a feature of a PAAS that we're building.
- sagichmal 5y ago> I think you're reading a critique of Prometheus that isn't really present in what we're writing. I'm observing that you've used pull-based, horizontally-scaled tools to build a push-based, vertically-scaled telemetry infrastructure. It can be made to work, sure, but the solution is an impedance mismatch to the problem.
- sevagh 5y agoI agree with you here. Using Prometheus, federated Prometheus, and Thanos on top of it for good measure, would probably get you better results without using a hodge-podge of non-Prometheus-compatible tools.