18 ms·
AWS CloudFormation now supports blue/green deployments for Amazon ECS
- ciguy 6y agoThe ECS default deployment model is quite frankly a disaster. I've written many Python/Go scripts over the years to wrangle it into a sensible form for CI/CD. Good to see that they're working on it, but I don't know why they don't fix the underlying paradigm instead of making it Cloudformation exclusive.
- justicezyx 6y agoCould you elaborate a bit on the problem of "ECS default deployment model"?
- ciguy 6y agoIt's just very rigid and not easy to extend. As soon as you hit a certain scale or need to do something slightly non-standard it's a mess.
- vazamb 6y agoI would love to know what the problem is. We do dozen of deployments every week with a ALB + ECS + Fargate setup. We upload a new container image, create a new task and launch as many tasks as desired (so if we want 2 containers running we launch 2, for a total of 4). ALB calls the /health endpoints on the new containers and if they pass the healthchecks it drains connections to the old containers and stops the tasks. This has worked seamlessly for a long time now without any downtime during deployments. EDIT: I should mentioned that we are using AWS CDK for all of this. All it does is register a new task as the default task for a service and ECS/ALB does the rest.
- Nk26 6y agoI would like to know the same, we have moved almost everything to fargate and ecs and have had zero issues.
- x3n0ph3n3 6y agoSame here -- I perform the exact same kind of deployment you mentioned using CloudFormation. My only grief is poor rollback detection / control.
- ciguy 6y agoThe default model works for most use cases, but it's not really flexible for non-standard cases and it fails completely at larger scale. If you have a few dozen containers taking 100,000 requests per second and you want to slow warm them with a segment of traffic for example, it's much easier to do on Kube than on ECS. ECS is also possible but it's just not as transparent or easy to work with.
- WatchDog 6y agoI find ECS particularly, ECS on EC2, to be really painful for small deployments. I just want a cluster that scales in and out to the amount of memory my tasks need, I'm not very concerned about CPU. Until capacity providers it was more or less impossible to do so without having a bunch of excess capacity provisioned. Even with capacity providers, I can't seem to get a cluster to scale in to 0 instances when no tasks are needed. I have one ECS service that requires an EBS volume mount, which means that when the task definition is updated, I need the service to stop, so that the new task can mount the same volume. This deployment model is essentially impossible without implementing a custom deployment strategy. Fargate makes things a bit easier, and now that you can use EFS volumes with it it might make more sense. But overall, everything in ECS just seems poorly designed, clunky and hacky, like many AWS products.
- friend-monoid 6y agoYou can’t do EFS on Fargate in CloudFormation yet though.
- scarface74 6y agoTrue. But you can do it with a custom resource. https://gist.github.com/guillaumesmo/4782e26500a3ac768888daab3c55b139 https://gist.github.com/guillaumesmo/4782e26500a3ac768888daa...
- WatchDog 6y agoYou also cant use capacity providers with cloudformation yet. Likely due to the fact, that you also cant delete capacity providers, nor change a cluster's default capacity provider.
- scarface74 6y agoThat’s not a Blue Green deployment.....
- ganstyles 6y agoI don't think they are saying it was, unless I'm misreading. They're talking about standard ECS deploys. Yo add my anecdata, I do lots of ECS deploys via terraform into production and it works pretty seamlessly.
- scarface74 6y agoI don’t have a problem with deploys with CF. But it only let’s you configure a minimum healthy percentage. Which is good enough if you only need to validate that an instance is in an acceptable state via a health check.
- sandGorgon 6y agoExactly this - I think internally there is a tussle on what is the strategic way forward. There are so many - Elastic Beanstalk, ECS, EKS, ECS-on-Fargate, EKS-on-Fargate..and of course the huge marketing push for Serverless. They could have the sensible way out and built EKS as the foundation of everything - makes total sense given the massive ecosystem around kubernetes. ECS and Fargate should be killed off. https://cdk8s.io/ https://cdk8s.io/ replaces Elastic Beanstalk ....but still basically runs on top of EKS. The pairing of CDK8S and EKS are fundamentally enough for all usecases that AWS basically sells.
- solidasparagus 6y agok8s is not for everyone. It has a high administration and complexity burden. It would be a mistake to make k8s a requirement for core AWS services. Fargate is a low-level compute platform - it is parallel to ec2 and lambda. It does not compete with kubernetes (as can be seen by the eks-on-fargate offering). There is probably some tension between ECS and k8s - AWS built a container orchestration platform based on what they think the world (and Amazon) needs and then k8s became madly popular. And it's not clear that AWS was wrong because k8s is essentially too complex for many use cases. It makes total sense that they would support both fully.
- harpratap 6y agoVanilla K8s is not meant to be directly used by everyone. Watch https://www.youtube.com/watch?v=ZqQTEdHVaCw https://www.youtube.com/watch?v=ZqQTEdHVaCw for insight on how K8s team expect the project to move forward. Frameworks like OpenDeis and Knative are meant to be used by developers. IMO K8s will become akin to Linux Kernel. Almost no one uses bare mainline Linux kernels, you choose a distro based on your needs. Companies like RedHat and Canonical will pop up and provide their own packaged "distros" of K8s and you will choose which philosophy best suits your needs.
- Androider 6y agoCDK8S doesn't replace Elastic Beanstalk, because the folks using EB are going to continue to use EB. If it ain't broke... ECS predates EKS, and is better integrated than EKS with their other services. If you want a painless container experience on AWS, ECS for Fargate is what'll you'll want to use today. So both existing folks who use ECS today as well as new customers are going to keep using ECS. Folks who like K8, perhaps from prior experience, are going to use EKS despite it having some rough corners. See, Amazon is perfectly happy to support everything under the sun as long as there are paying users. See their database offering for very much the same strategy, and it would be silly to say that they should drop DB x since they now offer DB y.
- bdcravens 6y agoWe haven't had any problems, but our workload is mostly async background work. Our CI pushes master when it passes, running a simple script to push to ECR with a git SHA tag, update the task to use the latest image, and then update the service to the latest task definition. It takes 2 lines of bash to update the task definition, and 2 lines to run the AWS commands.
- snypox 6y agoReally? At the last company we had like a 200 lines PowerShell script for all of that.
- scarface74 6y agoYou can do all of it by pushing the container to ECR, tagging it with your build number and running CloudFormation. You pass the build number in as a parameter and you specify your image as image: !subst “image:${Tag}”
- ciguy 6y agoYeah for the standard vanilla case it's ok. As soon as you hit a certain scale or need to do anything non-standard it becomes much harder to work with.
- tbrock 6y agoReally? We’ve been using it for 3 years and it works fine. Make sure you pay attention to min/max task settings. There are basically three ways to do it: Min less than max, max is 100. Max greater than min, min is 100. Some combination of the two. For the first you take down running tasks first and then backfill with new (green) tasks. For the second you add green tasks and once stable take down blue. Make sure that if you only have two tasks that min/max move in 50% increments: I.e you can’t scale up/down to 125%/75% but can to 150%/50%.
- grubbypaw 6y agoI've never understood CloudFormation's lag time in supporting new service functionality. This launched nearly 3 years ago. https://aws.amazon.com/blogs/compute/bluegreen-deployments-with-amazon-ecs/ https://aws.amazon.com/blogs/compute/bluegreen-deployments-w...
- x3n0ph3n3 6y agoBecause while "ready" for new features includes API support, it does not include CloudFormation support. The problem is entirely AWS politics and culture.
- jcims 6y agoWho builds the cloudformation support? I could understand not baking it in on an initial feature release but it should be there ’soon’. Same with Config.
- x3n0ph3n3 6y agoI've been told that each product team is supposed to add CloudFormation support themselves. For some reason, it's not treated as a high priority and often lands on the laps of the core CloudFormation team. It may turn out that the tools that CloudFormation team provides don't make integration easy, especially if the operation takes longer than 15 minutes (meaning they can't implement support via a single lambda invocation as an under-the-hood custom resource).
- scarface74 6y agoYou can also invoke a custom resource via SNS and then the SNS topic is subscribed to an API endpoint. But it’s hard to believe if the team responsible for the blue green deployment functionality developed an API endpoint to do it, they couldn’t just hand it to the CF team to call. At the end of the day that’s all CF does. Call APIs based on the different lifestyle events as far as how it actually creates resources.
- divbzero 6y agoI’m almost surprised this welcome feature was implemented at all, as ECS development appears to have slowed in favor of EKS (Kubernetes). Also too bad that it’s still CloudFormation.
- phamilton 6y agoIt looks like this uses a new, not yet documented, top level attribute in a template: Hook. At least, I haven't found any documentation for it other than the template they say to copy in the linked user guide for this feature. Definitely not supported in CDK.
- dkdk8283 6y agoWhat’s CDK in this context?
- mmansoor78 6y agoAWS Cloud Development Kit
- mcbain 6y agohttps://github.com/aws/aws-cdk/ https://github.com/aws/aws-cdk/
- laminatedsmore 6y agoThe list of restrictions in the user guide seems painful - cannot be used in the same template as nested stacks, cannot start a blue green deploy if other infrastructure would be modified at the same time, cannot import properties from other stacks, cannot export output properties. I was interested, but can't imagine using it with these restrictions.
- laminatedsmore 6y agoThinking about this further, the benefit of these restrictions is in avoiding conflicts where dependant infrastructure is updated in place before the blue green has completed. Lack of inputs and outputs still seems killer though.
- tbrock 6y agoI don’t think I’ve ever met someone that actually uses cloudformation unless a template was provided them to setup something like a lambda. Why would you when terraform exists?
- devonkim 6y agoTerraform doesn’t support rollbacks (handy for application roll-outs) and if your shop is heavily invested in AWS CF is a perfectly fine tool. I’ve maintained and started up tens of thousands of lines of both TF and CF and they both have their strengths and weaknesses.
- damagednoob 6y agoAren't rollbacks handled by the VCS that you check your Terraform files into? Sorry I'm not familiar with Cloudformation but that's how I would approach it with Terraform.
- takeda 6y agoCF can detect an issue while deploying, and then automatically go back to previous state (it is configurable, but that's the default behavior).
- tbrock 6y agoSure it does? Revert your code and apply. It’s not atomic but I’m guessing neither is CF. I’m not sure I’d use either of these tools to roll out new code though.
- devonkim 6y agoThat’s more of a GitOps style intervention requiring a source change and the workflow to revert a change could certainly be done but is by design not a first class construct in Terraform providers (not every provider supports every feature such as importing of resources). To respond to a Cloudwatch (heck, Prometheus, Grafana, ELK, etc) alarm saying your error rate went up because of a Route53 change it’s not out of the box with Terraform (would be a custom providers or null resources probably). And as a CLI application (granted, it’s run more like an RPC style architecture) there’s no obvious way to signal different failure levels and to respond to different failure modes of different resources. CF roll-backs are on by default and will revert changes. Sometimes it can fail and be a real pain, but it’s overall been more of a help for myself than harm.
- awsanswers 6y agoThe cloudformation team is small bc the people who run it are not fun to work with and not willing to bring on talent with their own opinions. It's a weird irony, dug in and overall detrimental. The only option teams that need CI/CD for their cloud resources & want to use cloudformation for most of it is to have a side chain process for handling resources that can't be expressed with Cloudformation. (not at all insurmountable but shouldn't be necessary)
- guitarbill 6y agoroughly how many is "small"? 10? 20?
- awsanswers 6y agoVery small - and to be clear this is not supposed to happen at Amazon or AWS it's "disagree and commit" vs "two pizza team". It'll suss out eventually but not soon. Start telling your teams to do what I said above and if anyone tells you "but we want to use one tool" tell them they must grow out of that.
- dragonwriter 6y ago> The only option teams that need CI/CD for their cloud resources & want to use cloudformation for most of it is to have a side chain process for handling resources that can't be expressed with Cloudformation. CF is extensible through Lambda (which AWS uses itself for the Serverless Application Model), which if you are going to use CF for this is probably what you bought to do to let you use it for everything, rather than having a “side-chain process”.
- tilolebo 6y agoI've been wondering: doesn't CDK depends on CloudFormation, as it basically converts code to CF files? (similar to Troposphere but with more languages supported). If that's the case,I don't see how CDK stands a chance against Terraform or Pulumi. And I'm not talking about multi cloud support, but just about the fact that an AWS-managed product, CloudFormation, keeps lagging behind for YEARS for such a core feature of a core AWS service.