5 ms·
Of course removing it reduced your CPU consumption. What did it do to p99 request latency? That was the trade-off being made by OOBGC. Does Github use SOA? The
by bluesnowmonkey 8y ago
Of course removing it reduced your CPU consumption. What did it do to p99 request latency? That was the trade-off being made by OOBGC.
Does Github use SOA? The more backend services that are involved in handling a user facing request, the more the latency of frontend requests will be dominated by the tail latency of backend services. So in distributed systems it makes much more sense to focus on those tail latencies. Who cares about average latency?
- hinkley 8y ago> Who cares about average latency? People who are bad at statistics. So most of us programmers. Seriously though, I don’t think many of us get around to acknowledging that microservices are distributed computing and then consider all of the limitations that go with it. Latency is dominated by the slowest thing you can’t start earlier. So in real distributed computing you see a degree of duplicate effort in order to improve latency. It reduces throughout but increases the utility of the system. When I took CS in college there was a proper class on distributed computing. When I got out I discovered not everybody had one of these (and I think in my program it was an elective. IIRC I chose it instead of kernel design). I had hoped that situation improved but I got tired of asking people disappointing questions about their education so I stopped asking.
- srean 8y ago> People who are bad at statistics. Heh! sadly very true. Many of those who do care about the p95s of an internal service often do it with a sense of chivalry that they have that 5% of the users back. The end user latency percentiles would be quite different from the service side latency, more so if its composed by a "Last_of" or a Max operator. At times I feel tempted to write a library that offers primitives to glue microservice calls such as fan_out, retry, timeout etc but one that can reason about the percentiles of the output from the percentiles of the input. BTW are statistical properties of latencies covered in distributed system courses ? I thought they were mostly about different kinds of guarantees, such as consensus, atleast-once, atmost-once etc etc.
- hinkley 8y agoHeck no, I got a little bit in a multivariate stats class but the rest I had to pick up in the field or a bit from books. I had a particularly illuminating moment with my ops guy trying to work out how often we could expect a drive to fail in a ten drive array and it was pretty shocking. My knowledge is pretty thin, but the fact that I even know what questions to ask puts me in a better spot to contribute than 4 out of 5 people in the room. That shouldn’t be how things are, but that’s where we are. Every new (old) technique is a bunch of naive people crossing their fingers and hoping things go better than last time.
- srean 8y agoI thought as much. Your handle has anything to do with https://books.google.co.in/books/about/Theoretical_Statistics.html?id=ppoujo-BInsC https://books.google.co.in/books/about/Theoretical_Statistic... ? very nice book BTW
- taf2 8y agoThere is definitely a trade off here... I noticed without OOBGC we did increase response times for the p99 and p95 requests however it was a few ms. So, for us the trade off of more CPU capacity and slightly slower in the p99 was worth the trade off. The cool thing about seeing this github post for me was to be reminded about OOBGC being a thing and that we should re-evaluate it after upgrading ruby.
- tenderlove 8y agoHi. Sorry, I should have included p99 and p95 info in the post (I wrote the post and shipped the patches). We didn't see any change on p99 and p95 requests, which is why we decided to keep it. I will include that info next time I make a post like this, sorry.