4 ms·
Netflix actually built their own metrics time series store called Atlas for similar reasons to Uber building M3DB (FOSDEM talk mentions hardware reduction and o
by roskilli 7y ago
Netflix actually built their own metrics time series store called Atlas for similar reasons to Uber building M3DB (FOSDEM talk mentions hardware reduction and oncall reduction), however open source Atlas only has an in-memory store component which was too expensive for Uber to run (since the dataset is in petabytes).
https://github.com/Netflix/atlas https://github.com/Netflix/atlas
- ckdarby 7y ago> which was too expensive for Uber to run (since the dataset is in petabytes). Ok, but I am fairly confident Netflix also is at that kind of scale. Netflix has a section on Atlas's documentation about how they get around this: https://github.com/Netflix/atlas/wiki/Overview#cost https://github.com/Netflix/atlas/wiki/Overview#cost They also did this nice video that outlines their entire operation including how they do rollups: https://www.youtube.com/watch?v=4RG2DUK03_0 https://www.youtube.com/watch?v=4RG2DUK03_0 This is how they do the rollup but keep their tails accurate to parts per million and the middle to be parts per hundred: https://github.com/tdunning/t-digest https://github.com/tdunning/t-digest
- roskilli 7y agoI want to first say, I have a great amount of respect for Netflix's engineering and for Atlas itself, it's great that it exists and is more accessible than other scalable in-memory TSDBs open sourced by large companies. A few of my thoughts on this, and this has come up before. Firstly Netflix self-identifies it is expensive to run an in-memory TSDB for metrics - for instance Roy's talk on Atlas mentions this as such[0] at the 37min mark of his Operations Engineering talk "It scales kind of efficiently. I'd love to say efficiently instead of efficiently-ish however that's hard to claim when my platform until this last quarter cost Netflix more than any other element of the cloud ecosystem ... Atlas and the associated telemetry costs Netflix 100s of thousands of dollars a week". At Uber M3 cost a significant amount to run as well at first and that is why M3DB was born to drive down that cost as much as it could and still provide a ton of instrumentation to engineers. Either way, giving engineers tons of room to instrument their code will result in a high cost no matter what since it will be viewed as a free lunch, that is why squeezing the economics on this matters since you want to provide as much instrumentation as possible at the lowest cost. Regarding your points about their documentation on cost: 1) Yes reducing cardinality by dropping node dimension on metrics, etc is possible to save cost - but also keeping things on disk is an alternate and great way to save cost too and keep the data at high fidelity. The challenge is making on disk lookup fast too, which with M3DB is what we were focused on doing. 2) Dropping replication of the data to a single replica is another way to save cost, however also comes with operational complexity as now you need to do backup/restore if you lose data and lose the ability to query that data in the meantime. This is why M3DB always is recommended (as per documentation) to run at RF=3 with quorum reads and writes so losing a single machine does not impact the availability of your operational monitoring and alerting platform. 3) Regarding rollups and tail solutions accurate, we always push for people to use histograms as that can be aggregated over any arbitrary time window and across time series. T-Digests are much more expensive to store raw and aggregate later. Bjorn talked about histograms, their use in Prometheus at FOSDEM[1] and why they're more desirable than t-digests or other similar aggregations. [0]: https://www.infoq.com/presentations/netflix-monitoring-system/ https://www.infoq.com/presentations/netflix-monitoring-syste... (video, quote is at 37minutes in) [1]: https://fosdem.org/2020/schedule/event/histograms/ https://fosdem.org/2020/schedule/event/histograms/ (slides and videos)
- pas 7y agoThanks for the FOSDEM link. I know the videos are out, but just looking at the schedule to find the interesting talks took more time than I wanted to spend on it. (The conference became so huge.) Maybe Bjorn's talk has the answers, but would you mind explaining how histograms are easy to aggregate? Don't you need either fixed buckets or raw data to produce a new histogram over a different dataset? (I know there are tricks to get great estimates, but naturally every re-aggregation would add larger and larger +/- intervals, no?)