2 ms·
The prometheus scaling problem is a real one. It stems from the "scrape everything" mantra that everyone in prometheus community follows, we should be rather op
by harpratap 6y ago
The prometheus scaling problem is a real one. It stems from the "scrape everything" mantra that everyone in prometheus community follows, we should be rather opting-in on which metrics are useful and drop everything by default. We need to go the distributed-tracing way, have smart agents near your metric source when gets everything locally and then decides what metrics are useful and finally ship them off for storage. We are kinda halfway there with OpenTelemetry but the adoption hasn't started yet. Very curious to see how you guys managed it.
- dima_vm 6y agoHey, we (Victoria Metrics) solved the scaling problem for you. We've yet to see a customer with data sizes our cluster version couldn't handle. I'd say "Scrape everything" mantra was born out of necessity -- engineer never knows what will be needed in the future. And "everything" didn't come for free either -- someone needed to expose that metric to scrape in the first place. If you drop most metrics by default -- it's much harder to add it back later and waiting for it to collect enough data. Users who take monitoring seriously, figured out long ago that it's much better to hoard everything that comes to mind from the beginning. [1] https://github.com/VictoriaMetrics/ https://github.com/VictoriaMetrics/