3 ms·
CI/CD would have have solved this 100%: > ... one of Knight’s technicians did not copy the new code to one of the eight SMARS computer servers. Knight did not
by movedx 3y ago
CI/CD would have have solved this 100%:
> ... one of Knight’s technicians did not copy the new code to one of the eight SMARS computer servers. Knight did not have a second technician review this deployment and no one at Knight realized that the Power Peg code had not been removed from the eighth server, nor the new RLP code added.
Read this part again:
> ... one of Knight’s technicians did not *copy the new code to one of the eight SMARS computer servers*.
Yes, of course a CI/CD pipeline can fail midway through and only partially deploy the code to a partial number of servers, but I doubt it. And even if that were the case, just off the top of my head I can guarantee an Ansible Playbook would have not only stopped the moment that particular transfer failed, the whole Playbook would have therefore failed, and none of the services would have been restarted (because that would be a final step that wouldn't be reached.)
This was due to human error and is the very reason CI/CD/automation is a thing.
> Knight did not have a second technician review this deployment and no one at Knight realized that the Power Peg code had not been removed from the eighth server
CI/CD would have have solved this 100%. A "Pull Request" made against a repository of Ansible code (or whatever you flavour is) would have *PREVENTED* the first technician from ever being able to merge the code into master/main (because you have master/main protected right... right?), completely preventing the entire process from ever rolling out without a review, which would have hopefully caught the misaligned configuration.
DevOps, which is mostly underpinned by CI/CD, would have solved this 100%. I'm very certain of this.
- SoftTalker 3y agoAnsible in my experience will stop trying to run subsequent tasks on a server once one of them fails, but it will go ahead with other servers that match the inventory pattern. So it very well could have successfully updated 7 out of 8 hosts. Maybe there is a switch that will stop everything if any task on any host fails but it's not the default behavior. At least it would have logged an error that hopefully would have been looked at.
- BoorishBears 3y agoI think this is an example of hindsight not always being 20/20 If you replace each step of the post mortem with a CI/CD based alternative, you miss out on the fact CI/CD trivializes designs where this wouldn't have happened. The "easy default" wouldn't be to run a play against 8 hosts manually in your terminal, it'd be run a playbook with them all baked in, and that would fail correctly by default: https://docs.ansible.com/ansible/latest/playbook_guide/playbooks_strategies.html https://docs.ansible.com/ansible/latest/playbook_guide/playb... The key here is CI/CD makes it so its actually less work to run that one play than it is to shoot yourself in the foot with 8 separate invocations. Even in the fact of incompetence/laziness/oversight, the general framework makes the right choice
- dfinninger 3y agoThe easy default here would be to have a runbook that executed on a particular inventory group. The “linear” execution strategy is the default (which you linked to). By default, if there is an error on one host it will continue executing on all other hosts. You need to set a flag to stop executing on all hosts[1]. The parent process would not be notified of any failures until the end of the run, unless you supplied a custom callback plugin[2]. [1] https://docs.ansible.com/ansible/2.8/user_guide/playbooks_error_handling.html#aborting-the-play https://docs.ansible.com/ansible/2.8/user_guide/playbooks_er... [2] https://docs.ansible.com/ansible/latest/plugins/callback.html https://docs.ansible.com/ansible/latest/plugins/callback.htm...
- BoorishBears 3y agoThe problem was a human forgot to run a step and no one noticed: The playbook would have failed and the server wouldn't have been online to make orders. If you read the article, the other servers were fine and did not contribute to the issue.
- dfinninger 3y agoYeah, it’s a runbook config: https://docs.ansible.com/ansible/2.8/user_guide/playbooks_error_handling.html#aborting-the-play https://docs.ansible.com/ansible/2.8/user_guide/playbooks_er...
- dtech 3y ago> So it very well could have successfully updated 7 out of 8 hosts. The problem was that the feature flag was manually enabled on the host with old code. Presumably with automated deployment the feature flag would never have been toggled if the deployment failed, either because the deployment didn't get that far or because the human spotted the failed deployment.
- piyh 3y agoTree, meet forest.
- Lutger 3y agoIt is quite likely this would have been solved by a good automated deployment process. However, it is also quite likely that at some point a human error would creep into either the automated deployment process itself, or be 100% correctly deployed into production. At that point, if the error is as serious, Knight would still gave gone bankrupt since they had no way to mitigate these failure conditions. Being 100% free of bugs is just not a viable way to end up with safe systems.
- movedx 3y agoAll processes are fallible, but some are less so than others because we've automated a large part of them. DevOps, and CI/CD, would have prevented Knight's issues 100%.