4 ms·
I was at Etsy when we made statsd, and I'm currently working on fixing the same problems at Stripe. This is super interesting and relevant to me personally, tha
by mcfunley 11y ago
I was at Etsy when we made statsd, and I'm currently working on fixing the same problems at Stripe. This is super interesting and relevant to me personally, thanks for writing it all down!
There are some details of how Etsy uses statsd that are not well-communicated. Etsy samples metrics aggressively to limit the amount of total traffic. And they monitor the packet error rates on the statsd boxes like hawks to keep the loss rate in check. Back when I was working there, if you added a high-volume counter without sampling it, the alarms would sound and you'd have an ops person tapping your shoulder pretty quickly. If you use statsd and skip either of these steps, the 40% loss that github experienced is what you get.
AIUI Etsy's moved to a consistent hashing scheme that's at least vaguely similar to this.
Node was not then, nor is now Etsy's area of expertise. We were going through an adolescent "let's just use every language" phase when we built statsd. I think the problems outlined here are solid supporting evidence that you should use a smallish set of tools and master them (a point of view which is very on-brand for Etsy engineering as it exists today).
- bbrazil 11y agoHave you considered pushing the aggregation into the applications, rather than doing it across the network? Having whatever library is sending data to statsd, instead keep counters/gauges in memory and then expose that on a regular basis would greatly reduce the data volumes involved as it's O(timeseries*frequency) rather than O(events). This is the approach we take with Prometheus, and based on the statsd setups of some people who've come talking to us there's scope for a reduction in network load of at least an order of magnitude without having to do any downsampling.
- mcfunley 11y agoYeah this has come up and it's reasonable, although it's a tricky/laborious migration in practice given a wide variety of things emitting stats. The statsd design choices here are mostly explained by the fact that Etsy uses it to collect from PHP. PHP doesn't afford a great way to aggregate in the client. (These are design choices that serve PHP well systemically, although it's limiting here.)
- bbrazil 11y agoI hadn't realised the PHP link, things make more sense now. You're getting into IPC then, which is a fun topic alright e.g. https://github.com/prometheus/client_ruby/issues/9 https://github.com/prometheus/client_ruby/issues/9 and https://github.com/prometheus/client_python/issues/30 https://github.com/prometheus/client_python/issues/30
- jrv 11y agoThis is how I scaled StatsD at SoundCloud (without having to change client code) before we switched to Prometheus: http://stackoverflow.com/questions/12871642/scaling-statsd-with-multiple-servers http://stackoverflow.com/questions/12871642/scaling-statsd-w... (the first answer)
- _wmd 11y agoNot suggesting it's a good idea, but the PHP standard library gives you enough tools (shmop) to allow a straight port of something like https://github.com/schmichael/mmstats https://github.com/schmichael/mmstats
- sciurus 11y agoIsn't the answer to the problem of network load to run statsd locally on each server? I thought that was how it was normally deployed. Then you can have the local statsd write to graphite/carbon directly, or to a second layer of statsd if you want to do additional aggregation.
- bbrazil 11y agoThat'd still have you going through the kernel on every event to handle the UDP packet, keeping in userspace is more efficient (on the order of 10ns of CPU).
- deleted 11y ago[deleted]