3 ms·
Unlike sibling commenters who just read about "thundering herd" problems on the Internet, as someone who spent significant time in SRE roles, I agree with you a
by solatic 2mo ago
Unlike sibling commenters who just read about "thundering herd" problems on the Internet, as someone who spent significant time in SRE roles, I agree with you as a matter of what the default approach should be. If you have a small cluster and no more than a handful of services... what thundering herd problem is there supposed to be, exactly, with so few "cattle" in the "herd"? Meanwhile, there are serious benefits, as you describe.
Large clusters with dozens of services and traces that go several services deep, with each service owned by a different team, are a whole 'nother ballgame, especially when overall production uptime is owned by an SRE team and not by the developer teams who wrote each of those services. And even in this scenario, you're not necessarily wrong; the risk attached to the cascading failure is domain-specific and may be acceptable.