4 ms·
You need better observability or tooling then. Operation teams (and automation) have usually a primary mandate of availability above all, not to investigate an
by chronid 3y ago
You need better observability or tooling then.
Operation teams (and automation) have usually a primary mandate of availability above all, not to investigate any possible failure.
- ilyt 3y agoYou achieve availability by redundancy, not by running around with shotgun murdering your cattle.
- chronid 3y agoDefinitely have multiple replicas, and recreate misbehaving ones, saving logs and data for later analysis. If you can't have that (and budget allows) keep unhealthy replicas alive and pull them off the load balancers. In my experience option one works best, and makes option two redundant. Unknown unknowns are a thing, but that's true even when the unhealthy replicas are kept around.