6 ms·
Instead of fixing something RIGHT NOW, meaning adding another commit to your build, why aren't you instead rolling back to a known good commit? Image is alread
by salamander014 6y ago
Instead of fixing something RIGHT NOW, meaning adding another commit to your build, why aren't you instead rolling back to a known good commit?
Image is already built.
C/I already certified it.
The RIGHT NOW fix is just a rollback and deploy. Which takes less time than verifying new code in any situation. I know you don't want to hear it but really, if you need a RIGHT NOW fix that isn't a rollback you need to look at how you got there in the first place. These systems are literally designed around never needing a RIGHT NOW fix again. Blue/Green, canary, auto deploys, rollbacks. Properly designed container infrastructure takes the guesswork and stress out of deploying. Period. Fact. If yours doesn't, it's not set up correctly.
- tleasure 6y agoSometimes you have to quickly roll forward.
- Raidion 6y agoThose sometimes should be super rare, and you should build testing infrastructure to prevent that from needing to happen. When you release something, you move traffic from the alb over to the new instance, if you have an issue, just move it back. If you are deploying breaking changes and don't provide yourself a stable upgrade and downgrade path, yea, you're gonna have trouble.
- tantalor 6y agoPlease do not.
- gregorymfoster 6y agoIncidents _requiring_ rolling forward are extremely rare. In the cases you have to, just build the image and deploy to your cluster with a high max-surge configured. If you image has correct caching, rebuilding it shouldn't take much time. Most of your time is likely spent in CI and rolling deployments, both of which you can manually skip.
- PragmaticPulp 6y ago> If you image has correct caching This is the hangup for most CI/CD systems with containers. Typical configurations (e.g. Gitlab basic setup) don't leverage any caching, so every container is built 100% from scratch, every time. Adjusting the system to properly utilize caching and ordering your container builds in a way that the most volatile steps are as late as possible in the build will massively speed up container builds.
- tashoecraft 6y agoI’ve learned to take the time and go through the normal deploy steps for any hot fix. More often then not, rushing the steps leads to longer outages, missing the actual bug, creating a new bug, etc. Don’t cowboy it, deploy properly and you’ll be more relaxed in the long term.
- tleasure 6y agoYeah, I should have been more clear. I'm 100% for using normal deploy steps and I'm not recommending cowboy-ing updates in a container. He was asking about using non-containerized infr though. If you can commit a code hotfix and quickly deploy the code package, you can roll forward without the slow container build/deploy.
- greggman3 6y ago> Properly designed container infrastructure takes the guesswork and stress out of deploying. Period. Fact. If yours doesn't, it's not set up correctly. Hmmm, I don't have the experience to know if it's setup correctly or not. All I can do is watch it fail and then learn from my mistakes. Is there a container "framework" that out of the box gives me all of " Blue/Green, canary, auto deploys, rollbacks..." so I don't have to guess if I'm doing it right?
- ohthehugemanate 6y agoI hate to say it, but yeah that's kind of the point of kubernetes deployments (https://kubernetes.io/docs/concepts/workloads/controllers/deployment/ https://kubernetes.io/docs/concepts/workloads/controllers/de...). Or openshift for more UI and "out of the box" experience. You deal with all the headache of making your app stateless with a predictable API so that you can reap the benefits of a system like k8s, which can automatically manage all of it for you. Similarly i'm a bit confused by your comment about SSH dying... in k8s you configure a readiness/liveness probe and behavior when the probe starts to fail. If SSH is an important thing for a given container, maybe the "liveness" probe is the command "ps aux |grep sshd". Then if it dies, the container can be pruned automatically.
- erikcw 6y agoWe’ve been using Convox[0] for the last 2 years. I’ve been pretty happy with how simple it is to work with. We’re still on version 2 which uses AWS ECS or Fargate. Version 3 has migrated to k8s and is provider agnostic. We just haven’t had the bandwidth to upgrade yet. [0] https://convox.com/ https://convox.com/
- tehalex 6y agoWe are using Convox v2 too and are happy with it, but I'm hesitant to do the upgrade to introduce the complexity of kubernetes to our devs and if convox the right abstraction on top of kubernetes when there which is already a pile of abstractions in k8s itself (and so many other tools to choose from in the k8s universe). https://github.com/aws/copilot-cli https://github.com/aws/copilot-cli isn't ready for our use cases, but is more or less convox v2 built by AWS.
- jejeyyy77 6y agoYou sound like you might be emotionally invested. Fact. Period. End of story.
- milofeynman 6y agoMy guess is that they are writing code (migrations...) Without thought given to rollback.
- freedomben 6y agoI very much disagree. If your bad deploy included migrations that can't be reversed (drop an old table for example) then rolling back just gave you two problems. As long as the dev wrote (and tested!) two-way migrations and they are possible, then yes you are correct.
- jasonhansel 6y agoIMHO the answer here is to just never do irreversible migrations. (In other words: just leave old tables/columns around--if not indefinitely, then for a few weeks after they stop being used.)
- freedomben 6y agoI have done that in the past and I agree, it's a decent solution. The risk is that the old stuff never gets cleaned up and before long your DB is full of all kinds of cruft and nobody knows which part is needed and which isn't. I saw a heinous bug once because a newer person was using a column that had stopped getting updated years earlier (because it was replaced). This person saw the column and it was exactly what they needed. Then customers started getting billed on old accounts that they had either closed or changed with us. They were really mad.
- jlangenauer 6y agoMy solution to this is to rename that table/column to “table_deprecated” or “column_deprecated”. This has the nice property of being reversible, causes nice visible errors if something unanticipated is actually using the column/table, makes it obvious that it shouldn’t be used (well, one would hope), and makes it easy to find for permanent deletion later.
- freedomben 6y agoThat's a great idea! Although I did work with a person once who would use anything that looked helpful regardless of "deprecated" being all over the place (Eclipse would even yell at the method calls but he didn't care). Granted that's Java and at that point in time APIs were being deprecated without workable replacements, so it's hard to fault him.
- deleted 6y ago[deleted]
- rurp 6y agoThis seems overly absolute to me. What about all of the cases where the bug wasn't caused by a recent commit? Some cases of this I've seen are: * Time bomb bugs. Code gets committed that works fine until some future condition happens, such as a particular subset of dates that aren't handled properly. * Efficiency issues. Some code could might function properly and work fine with low amounts of data, but hit a wall when it has to handle loads beyond a particular size. * Bugs in code that just hadn't received much traffic yet. A feature having a bug that only affects 0.1% of people using it might not be discovered until the feature gains traction down the line.
- Azkar 6y agoHaving dealt with all of these in production, I can tell you the strategies I've used to combat these things: 1. Solid code reviews. Anyone of our developers can halt a code review for any reason. We require 3 approvers on each review. Sensitive areas require reviews from people familiar in that area. We also have tooling that allows us to generate amounts of test data in dev that is similar to prod loads. This helps us catch a lot of time bombs. 2. Feature toggles to decouple deploy of code from release of code. This allows us to test our code in production before turning it on for customers. It also allows us to slowly rollout a feature and watch how the code behaves. This also gives us a kill switch to turn off the code if it is bad. 3. An incredibly robust testing pipeline. It takes about 50 minutes from commit to production deployment. We can also deploy previous containers very quickly for situations that require it. This doesn't solve all of our problems. Some changes cannot go behind feature toggles (DB migrations, dependency upgrades, etc). But we do pay a lot of attention to design and rollout plans for database migration changes and such. All of these things come at an extra cost to us, but it allows us to move quickly when we need to. But we're in a lot better place than we were when we were trying to do weekly releases. We have a good mix of team experience (sr vs jr) - and have a lot of discipline in our software engineering practices. We still have problems like I said, but these strategies have greatly improved our ability to deliver software.
- JoshuaDavid 6y agoOut of curiosity, how many devs does your org have? I think a lot of the disagreements here come from people at orgs with 3 developers talking to people at orgs with 30,000.