3 ms·
1) you cannot have a runbook for everything, and even if you have a runbook you the best you could have found in this case is that something weird was happening
by amessina1 6y ago
1) you cannot have a runbook for everything, and even if you have a runbook you the best you could have found in this case is that something weird was happening in the VM. The setting had an insanely big value but it was accepted by the kernel, so you would assume it was a valid one.
2) the customer provided a huge amount of details, but it is usually very hard to explain what did you change from the base image. Most customers might not be willing to provide the full spec of their running system, as they might contain information they don't want to disclose. It is easier when the issue starts right after a change has been made, but this was not the case.
- rantwasp 6y agoi’m not saying a runbook will catch everything. but it will give you a chance to solve the problem quicker.
- amessina1 6y agoI totally agree, but keep in mind that it is an iterative process: you have a case, you apply the runbook/playbook you have, if they are not enough you use your skill/knowledge to solve the case, then you update the playbooks.
- rantwasp 6y agoiteration is my middle name :) the point I was trying to make is you should be prepared - if only for the 95% of the cases that the runbook can solve. you also don't start with a complete runbook - you build it as you operate the service.