3 ms·
Whether or not this theory is true, can be measured. A few important numbers come to mind. - do you get more operations tickets right after a production deploy
by objectified 12y ago
Whether or not this theory is true, can be measured. A few important numbers come to mind.
- do you get more operations tickets right after a production deployment?
- do your call centers get an increased calls/hour rate after a production deployment?
- are there often noticeable anomalies in system resource usage that seem to be directly related to your deployment cycle?
- do your monitoring tools show a higher rate of warnings/criticals right after a production deployment?
Whether or not production deployments introduce more risk is probably largely subjective. How well do you test, what number of changes are inside a deployment, how collaborative are your operations/development teams, are your technical teams understaffed, and so on.
The advantages of a freeze (defined as "no production deployments during a certain time window") that I see, are at least the following.
- from my own experience while working in operations, I remember very well the many sleepless nights that somehow always seemed to be the immediate consequence of new (byte)code running in production
- it gives operations a break; they often work at ungodly hours the whole year, and getting some rest is very much needed
- it gives a moment to stand still and reflect; work on internal tooling, putting some structure into ad hoc things that sneaked in, etcetera.
Furthermore, I don't agree with the philosophy that "everything is always broken". Sure: disks and power supplies break all the time, even load balancers break, security patches need to be applied, and so on. But these are things that are part of day to day operations, and most operations engineers know how to do them. It's usually controllable. Unlike a bug in some newly introduced code by one developer that causes a stack overflow every time one certain application flow is being hit. That requires a different kind of discipline to solve.
I think it's a little dangerous to generalize these things without having actual numbers to back them up; before you know it, your operations team won't have any excuse to just sit and play Quake for a week. In most cases, that's a joke.
- dasil003 12y agoI agree the OA is a bit disingenuous about acknowledging the risk of deployments, but the four metrics you came up with also don't tell the whole story. Of course there will be more breakage after a deployment, the real question isn't whether that's true or not, it's whether subsequent deployments will becomes even more risky by withholding earlier deployments.
- Retric 12y agoNot all risks are equivelent. There is a reason planned outages are scheduled for ~2AM local time not ~2PM local time. Plenty of companies have dealt with 2h windows where if the site is down they fail. Some have even been down during that time period and ended up laying off everyone.