8 ms·
This is very common and just because it doesn't match your use case doesn't mean some businesses don't need a hot fix "right now." If you work in 24/7 ecommerc
by one2know 6y ago
This is very common and just because it doesn't match your use case doesn't mean some businesses don't need a hot fix "right now." If you work in 24/7 ecommerce and your site is producing $60k per hour and there is a network failure that breaks something you need a hot fix right now, otherwise your 3 hour code review, build, q/a, deploy pipeline will cost the company $180,000.
- imdsm 6y ago> otherwise your 3 hour code review, build, q/a, deploy pipeline will cost the company $180,000 Taking code review away, because we all know we're not going to sit around waiting for a review while a patch needs to go out urgently, if your build, test, and deploy pipelines take three hours then you have some serious problems that you need to address, and containers aren't it. There are methods for handling hotfixes/patches to production quickly that work well in high volume sales website setups.
- pdimitar 6y agoPlease do tell. Our containers never take less than 10 minutes to deploy (a single one).
- asdfaoeu 6y agoWell 10 minutes is way less than 3 hours.
- pdimitar 6y agoCertainly. But it does add up and it does kill motivation for rapid iteration. :(
- Guthur 6y agoand then we're back to the original statement, if you're rapidly iterating hot fixes you have serious problems and likely doing it wrong.
- jonpurdy 6y ago10 minutes to deploy to production should be fine, but rapid iteration shouldn't be happening in production. It sounds more like the complaint is related to running a similar setup in dev, and taking 10 minutes to see changes during development (which is understandably too long).
- pdimitar 6y agoYep, that's what I meant. I have no choice but to run a local k8s cluster or else I can't test.
- owenmarshall 6y agoI have app stacks that deploy a dozen containers in seconds because they are stateless and close to "functional" – just transformations over inputs. I have app stacks that deploy a dozen containers over an hour because the orchestration takes time: signal the old containers to drain, pause for an app with a very long initialization time to settle, gradually roll traffic to the new one to let caches warm, and then repeat. In both of these cases, deployment is a function of the application. There's nothing infrastructural that puts a time floor on things.
- pdimitar 6y agoSure, not contending that. It's just that in my memories blue/green deployments still took less time, although I can't say how much.
- ryanjshaw 6y agoNot a serious reply, but it may interesting: One million containers over 3500 nodes in 2 minutes: https://channel9.msdn.com/Events/Ignite/Microsoft-Ignite-Orlando-2017/BRK2190 https://channel9.msdn.com/Events/Ignite/Microsoft-Ignite-Orl...
- mumblemumble 6y ago> Taking code review away, because we all know we're not going to sit around waiting for a review while a patch needs to go out urgently I worked at a trading firm where you could add a zero to that cost for a 3-hour outage, and I can tell you that the one thing we would absolutely never skimp on was code review. Because the cost of a bad "fix" that actually makes things worse has the potential to be greater still, and because humans are most likely to make silly mistakes when they're working under intense pressure. What we would do instead is slightly intensify the code review process by pair programming the hotfix, and ensuring that a third developer who was familiar with the system in question was standing by to follow up with an immediate review.
- codemonkey-zeta 6y agoI really like that approach. It also takes pressure off the tech lead or whoever is implementing the fix, and transforms the patch into a full-team responsibility. I bet this sort of behavior makes for strong and effective teams.
- deleted 6y ago[deleted]
- allannienhuis 6y agoPair programming anything as urgent as a hotfix works really well. It takes some pressure off of the developer working on it and turns it into a team event. We will even sometimes keep the video call open until deployment is done and production validation is complete - the devops guys get the information they need, someone else is keeping an eye on the checklist and calling out items if necessary, etc.
- mumblemumble 6y agoOh yeah, absolutely. I was happiest when I also had the attending ops person looking over my shoulder while I worked on the fix.
- lovehashbrowns 6y agoA lot of businesses that operate 24/7 run on containers quite well. When I had to so these sort of quick hot fix things for a startup, almost all of the issues were caused by lack of testing. Testing lacked because there wasn't an easy way to constantly make sure dev staging and prod are absolutely the same. Same infrastructure, same code, same packages, etc. That's an easier solve with docker containers. And testing, including UI testing, can be integrated much more easily with ci/CD tools and docker containers that have code which goes by commit hashes and which ensure every package down to the version is controlled across environments.
- quickthrower2 6y agoSomeone fucking directly with prod might also cost $180000.
- FpUser 6y agoOne will cost and another might, see the difference
- nostrebored 6y agoTime to deploy is a known value. The impact potential of having a workflow where ssh'ing into production is even possible can cost you buckets as: 1. You're messing with production, obviously. 2. Your infrastructure isn't stateless. 3. Your infrastructure is likely not HA. 4. You likely don't have canaries in place to mitigate the impact of bad production deployments. And the impact of all of these on direct revenue/productivity can be immense. SSH into prod is a crutch.
- kugla 6y agoIn the long run "fucking with production" as well <i>will</i> cost a fortune.
- iso947 6y agoCitation needed. I’ve “fucked” with many prod systems over the last 15 years and not caused outages, or extended them.
- quickthrower2 6y agoThey are both a might. Both are a probability distribution.
- jpz 6y agoThose costs need to be amortised annually. There are other costs from outages which occur infrequently but are highly costly because developers/support are manually touching the servers. The cost really should be analysed as what is the median cost on revenue (or profit) as a percentage point. One off pricing is pretty meaningless.
- malka 6y agoyou rollback and then make a proper fix.