6 ms·
roll back (step back), is an inherited from waterfall anti-pattern. Now we should only march forward with small, on demand releases, this way we will know exac
by rinchik 7y ago
roll back (step back), is an inherited from waterfall anti-pattern.
Now we should only march forward with small, on demand releases, this way we will know exactly where the issue is and will be able to fix it forward quickly.
Rollbacks were a strategy with monthly (or even quarterly [insane huh?]), giant, stinky, release dumps, knowing there is no way we could quickly identify and deploy the fix. aka lets throw production 3 months back and take another 2 month for figuring out there the issue that happened during last release is.
And to finally answer your question: we never roll back. We always march forward.
- ISL 7y agoUntil the fix is identified, can't one 'march forward' by rolling back recent changes to a known-good state?
- rinchik 7y agoNope, you are "pretending" that a step back is a step forward (which is not true). I never had to roll back anything within last couple years and very happy about that. Also note, that roll back was a valid strategy back in the day and still can be useful tool in your garage of tools. It can be useful when, for example, dealing with complex legacy systems that were created decades ago, or complex systems developed by outsourced development teams. You'll know when roll back is useful when you see it.
- nprateem 7y ago> Nope, you are "pretending" that a step back is a step forward (which is not true) If the business is losing $BIG_BUCKS per minute of downtime, it is most definitely a step forward.
- mooreds 7y agoCongrats, that sounds awesome. So when a bug affects prod, do you: * Find the code that is affected * Write a test * Have it go through ci/cd * Deploy to prod. Or is there a different way of deploying a big priority bugfix to production than normal deploys?
- rinchik 7y agoLas 2 steps are merged (ci/cd is part of the deployment to prod, but generally yes. Also worth to note, priority bug fix is not really about pipelines it's more about the ability to dynamically reallocate resources. Depending on the complexity of the affected area we should be able allocate as many devs as it is useful to fixing it. (similar to "Fast Lane" in Kanban)
- kostarelo 7y agoI totally get your point but this approach seems a bit dogmatic to me. Even with modern CI/CD techniques and agile methodologies in place, rolling back could still be the best choice. For example, always marching forward means that any time an issue is coming up, certain resources must be allocated in tackling the issue. That can't always be the case. Smaller and more frequent releases are preferable and most of the time a single line change will fix the issue, but other times rolling back may be the best option.
- gfodor 7y agoShorter releases can help with reducing the difficulty of immediately addressing problems, but it's an error to equate the reduction of risk as an elimination of risk. There are always going to be failure modes that require extensive time to diagnose and debug, even with small changes being made. Additionally, you want that diagnostic phase to happen without time pressure. If you do not have a sane rollback mechanism to use in those scenarios, you are doing a disservice to your users and your team. Your users suffer, because the outage or breakage will last as long as it takes for you to address the underlying issue directly, instead of just rolling back to restore service. They will be forced to hear frustrating things like "we're working on it", since you don't know what's wrong yet, when instead you could have just rolled back before most users even noticed there was a problem. And, more importantly, your team will suffer greatly, because they will be forced to work under pressure when an incident like this arises. And, worse, they will also 'learn' that accidentally pushing breaking changes to production results in an extremely unpleasant and toxic situation for everyone, leading to systemic fear-of-deploys and undermining a blameless culture. So you should have a rollback mechanism that is solid, tested, and easy to use for scenarios where a non-trivial regression or outage arises in production, even if you are doing continuous delivery of small patches.
- rinchik 7y agoWell, highly unlikely that such an error will arise. If it does and you know that you actually need to roll back then most likely something else is wrong. But again, I also saw the other comment about how "Dogmatic" my approach is. I wouldn't say it's dogmatic, idealistic - yes. But not dogmatic. There is a place and time for anything and roll back can STILL be useful when you don't trust the system nor the code base (as I pointed in my other comment, rollbacks are useful with legacy systems and systems that you have to maintain that were build by outsourced teams). Well also roll backs is the first thing you think of when you join a large company as a director of engineering to support systems you never touched before.
- gfodor 7y agoYou didn't really address my points. Its hard to quantify just how "highly unlikely" a failure is, but its your job as a systems designer to build systems that are robust under a wide variety of unlikely failure scenarios. Not having rollbacks results in a system that is extremely problematic in those unlikely scenarios where a quick fix cannot be immediately addressed. Not to mention, rolling forward under such a regime has its own unique risk: since the 'fix' was made under pressure since there was no alternative to restore service quickly, it's often the case that simple human errors get introduced when rolling forward. I've seen it a million times. In my experience, a healthy incident response process has a fork in the decision tree at the very top: do we roll back, or do we attempt to fix live? And in the latter case, we time box how long we're willing to spend, and defer to rolling back for all but the most trivial, obvious fixes. Even if you don't use rollback often, having that top level fork is a release valve for all of the toxic implications I mentioned in the scenario where you do actually need it. Even if you have several dozen incidents happen where you didn't need it a black swan event will eventually show up -- and that event will be the one that will have the lasting impact on your company's public perception and the morale of your team.