3 ms·
Restart should be a very last emergency step, as if it works, a restart often might wipe out evidence of why. So hopefully it's not done often.
by b112 10d ago
Restart should be a very last emergency step, as if it works, a restart often might wipe out evidence of why.
So hopefully it's not done often.
- CoffeeOnWrite 10d agoI wouldn't say so, rather you need to balance recovery time and evidence preservation. A good incident manager will give the service owning team a chance or two to debug, but not let them fall into the trap of needing to understand the problem fully before attempt a clumsy potential fix. And of course will take into account the total business impact of the ongoing disruption and the known and unknown risks of the proposed clumsy fix (it could make things worse).
- b112 10d agoYes, that's why it's the last emergency step. We're not disagreeing.
- ocdtrekkie 10d agoRebooting is the first step: If it fixes it you don't have a problem. If it doesn't, you know more about the problem. It's a joke, but like, after nearly two decades of engineering I have something break on me, due to updates. I figure the updates broke it, call the vendor, and they go... did you try rebooting it again? Rebooting it a second time fixed it.
- wookmaster 9d agoThe requirement is to get customers out of impact as #1 priority. If there's suspicions around memory/thread states a restart makes a lot of sense. Digging through logs and flight records takes a lot of time, customers are losing business in that time. If you're afraid to restart your service you need to work on your telemetry.