3 ms·
You've observed something important: deployment is a frequent source of errors. It is a discipline unto itself which deserves understanding, study, coding, and
by aamar 15y ago
You've observed something important: deployment is a frequent source of errors. It is a discipline unto itself which deserves understanding, study, coding, and infrastructure. Here is my system:
First, here are the specific types of solutions, from best/most-difficult to least-efficient/most-achievable:
1. Automation, e.g. have the date automatically set by the deployment your deployment script. Make your scripts be smart, so they can set things correctly.
2. Poka-yoke (http://en.wikipedia.org/wiki/Poka-yoke http://en.wikipedia.org/wiki/Poka-yoke): e.g. have the deploy script refuse to push out code if a configuration variable is unchanged, unless there is a comment on that config line overriding the check.
3. Checklist: Write down a set of procedures/checks to do when deploying. Make a copy of this list on deploy, and check off each item as you do it. (See also: http://www.newyorker.com/reporting/2007/12/10/071210fa_fact_gawande http://www.newyorker.com/reporting/2007/12/10/071210fa_fact_...)
You'll note that all of those are geared at reducing reliance on your memory, rather than improving your memory. Next, here is the overall process:
- ("5 whys") When you have a problem in deployment, write down what went wrong. Find 5 ways that problem could have been prevented.
- Address at least two of those with solutions from the set above. For example, you might automate part of the problem and add a poka-yoke in another script. Or if you're ambitious, you may be able to have two fully automated solutions, an automated deployment script and an automated test which checks that that automation is working. Not as good (but sometimes easiest) is to add it two different checklists, filled out by different people or at different times.
- Implement solutions even when something almost went wrong but didn't.
- Periodically review the overall process; refactor and otherwise improve it. In particular, push solutions up the chain, e.g. replace a checklist item with a much better automated solution.
Additional things to consider, specific to deployment:
- Deployment frameworks (e.g. Chef) can assist with automation.
- Deploying to a staging environment first can flush out many issues.
- Deploy first to a small subset of servers/users, following with deployment to all users X hours later.
- Get the advice and support of an experienced NetOps or sysadmin person.
Edit: shouldn't be so negative about checklists; they're often useful.
- thisrod 15y agoI use checklists, and they work. Scientists and engineers have a technique for complicated mathematics, which might work for editing sendmail.cf. Do it twice, and compare your answers. This works best when the "you" is plural, but repeating your own work the next day is better than nothing.
- Evgeny 15y agoThere's a lot of useful advice here, and I'll do my best to action on this ... The most important part will be to get enough discipline to apply the system - a challenge in itself!