16 ms·
Our nightmare on Amazon ECS
- hosh 10y agoI put something into production with ECS as well. I ran into the same missing components too -- lack of service discovery, and such. Kubernetes work a lot better. As it stands right now, I wouldn't take a gig involving putting ECS into production. Now if ECS 2.0 was really AWS hosted Kubernetes, I would be very interested in hearing about that...
- moondev 10y agoThat's exactly what GKE is on GCP and I love it.
- tantalic 10y agoGoogle Container Engine (GKE) is certainly the easiest way to setup a Kubernetes cluster. We have been running it for a couple of months now and couldn't be happier. If you're wanting to stick with AWS I have always heard great things about the work CoreOS has been doing in this space: https://github.com/coreos/coreos-kubernetes https://github.com/coreos/coreos-kubernetes.
- hosh 10y agoThat is what I keep hearing. With PetSets rolled out in 1.3, GKE is getting more competitive. At my current job (startup), we're probably going to move towards that.
- alex-mohr 10y agoIt's great to hear GKE is meeting your needs so well! (Yes, I work on it.) For 1.4, the Kubernetes Cluster Lifecycle and Ops SIGs are working on making the install and setup process much easier, including on AWS [1]. That won't magically turn it into Kubernetes as a Service, of course, but we hope it'll help users on other platforms. [1]: https://github.com/kubernetes/community/blob/master/sig-cluster-lifecycle/README.md https://github.com/kubernetes/community/blob/master/sig-clus...
- advisedwang 10y ago[off topic] Author, if you are reading this be aware that when viewed in a narrow browser window the sharing icons overlap the text, even though 40% of the screen is taken up by the right hand sidebar/empty space.
- maslam 10y agoThanks @advisedwang. We're looking into it.
- cddotdotslash 10y agoWhat I don't understand is why AWS squandered this opportunity. Given the popularity of Lambda, they clearly saw the market for completely managed services. They could have designed a platform where users upload containers and AWS runs them. No servers, no crazy settings, etc. Instead, they created this entire platform where you still have to run the entire EC2 infrastructure, there is no service discovery, etc. They essentially created a half baked Mesos or Kubernetes clone. I'm still shocked when I hear companies going "all in" on ECS.
- codemac 10y agoA good friend told me that he felt that Google Cloud and AWS had "a severe lack of imagination". Imagine what you could do if you didn't even assume a process model? All app state just resident in memory, but magically persisted? Who needs object storage, re-invent the pointer! We could have lived in the future, now it seems we're permanently wed to the past.
- mschuster91 10y ago> Imagine what you could do if you didn't even assume a process model? All app state just resident in memory, but magically persisted? Who needs object storage, re-invent the pointer! Take your usual Java, NodeJS or Ruby payload, enjoy your memory leaks eating up your space.
- count 10y agoWe're starting to get to the point where these giants can innovate like that. It wasn't 3 years ago that 'nobody serious' was 'trusting' cloud providers like AWS/GCE with anything important. This is still the very early days, as evidenced by the ridiculous growth numbers being posted YoY.
- gtaylor 10y ago> A good friend told me that he felt that Google Cloud and AWS had "a severe lack of imagination". I don't know about that. Google Container Engine (hosted Kubernetes) is actually pretty awesome and imaginative. It's feeling like GCP's niche is going to have a sizable containerization element. If you browse around their docs, you'll find that GKE/containers have started creeping into the examples for seemingly unrelated services. They're not just dipping a toe in. More generally, I feel like GCP's container strategy is just leagues ahead of AWS' at this point. While this article was thin on substance, ECS is definitely difficult to set up and maintain. If I'm going to go through all of that trouble, I might as well run my own Kubernetes or Mesos setup and not be locked into ECS.
- maslam 10y agoHN, I'm a co-founder at Appuri. Happy to answer questions! PS: We LOVE most AWS services like Amazon Redshift. Just not ECS ;)
- ntumlin 10y agoOff topic from the article but just wanted to let you know that I love the design of your blog.
- maslam 10y agoThank you!
- dstroot 10y agoDid you deploy K8S on AWS? If so can you add any details about how? Or are you using K8S elsewhere? I love AWS but planning on spinning up on GCE this weekend to play with K8S.
- tmacie 10y agoWe deployed K8S on AWS (I'm a dev at Appuri). Like Bilal mentioned we run pretty much everything on AWS, so it was an easy decision.
- lobster_johnson 10y agoThe kube-up.sh script can start a complete Kubernetes cluster on AWS in one fell swoop. It's pretty smooth.
- velkyk 10y agoI am ops at Appuri. We deployed k8s on AWS since we are using other services like Redshift and RDS inside the VPC, also happy with how EC2 works plus of course we have Reserved instances so we didn't look into GCP yet. We're running kube on CentOS 7, we bootstrap nodes using cloud-init (user-data) to setup k8s, which we then use to run everything else. I would love to give you more details, I might write blog post about our kube setup decisions later. Kelsey wrote nice manual for setting up k8s - https://github.com/kelseyhightower/kubernetes-the-hard-way https://github.com/kelseyhightower/kubernetes-the-hard-way which is definitely on my to-read list this weekend :)
- SteveWatson 10y agoArticle text is obscured by icons.
- maslam 10y ago@SteveWatson - thanks for reporting, should be fixed now.
- cyberferret 10y agoExcuse me while I pick myself up off the floor when I read "leaks environment variables"... What?? That is incredibly scary for use, as we just went through an audit process of our code on about 6 different web apps to ensure that all secrets were placed in environment variables on our Elastic Beanstalk configs and not in the main codebase... If this now results in LESS security (all our code is in private Git repositories) than before, then we have essentially taken a step backwards!
- dorfsmay 10y agoGithub private repos have been made public by mistake before. Got repos are cloned on dev laptops, do you enforce laptop encryption? The right thing to do is using some form of a vault.
- cyberferret 10y agoWe use BitBucket here, rather than Github - similar risks, I know, but we have predetermined repositories which are all set as private. 3 dev machines which are kept on premises at all times. Still not optimal as far as security goes, but it seems that he have roughly the same exposure if AWS leaks our keys and passwords to other third party trackers...
- Ixiaus 10y agoUse kms and dynamodb with key enveloping, or this tool: https://github.com/fugue/credstash https://github.com/fugue/credstash Don't initialize into env vars and don't store in repos, even private ones.
- cyberferret 10y agoThanks - sounds like a good solution. I will look into this in detail over the weekend.
- fletchowns 10y agoBe careful when modifying user access to a private BitBucket repository. Their autosuggest for the username input field will show all bitbucket users. Makes it incredibly easy to accidentally grant somebody outside of your organization access to a repository.
- huslage 10y agoEnvironment Variables are NEVER private. Please don't think that you can hide information in there as all of that information ends up in the process table which is public across the entire machine.
- nnutter 10y agoHow are they "public" across the entire machine?
- phil21 10y agoI suppose it depends on your definition of public. Any environment for a process will be accessible via /proc/<pid>/environ on a linux system. Of course other users cannot read these files, however in the case of something like a Docker image all processes likely share a username and this could be a risk (especially for a public webapp that one day may information leak/allow remote command execution). At least that's my immediate take on it.
- nathanboktae 10y agoAnd the machine (vm) runs in a private VPC. So it's private.
- zeroxfe 10y agoWait, what? How do environment variables end up in the public part of process table? There's no way one user can peek into the environment of another (without permissions, notwithstanding bugs) -- that's part of the design of Unix systems.
- Ixiaus 10y agoProbably want to use a secret management tool and just not initialize into environment variables... https://github.com/fugue/credstash https://github.com/fugue/credstash
- justicezyx 10y ago“No central config. ECS doesn't have a way to pass configuration to services (i.e. Docker containers) other than with environment variables. Great, how do I pass the same environment variable to every service?” Would packaging the configurations together with the docker image makes more sense? That enables more hermetic deployment.
- velkyk 10y agoDo you mean hard coding configs to docker image? I wouldn't support this, IMO this is worst case scenario setup :) Imagine you need to change single config value, for this you would need to update image, push, build, redeploy, this can take some time depending on your deployment. With k8s you do only `kubectl edit configmaps <name>`, restart pods that are using it and you are done. Also no need to creating per stage images...
- jbaviat 10y agoWe have been running Sqreen production on ECS since October 2015, and we have been pretty happy about it the whole time. Of course ECS was very minimal at the beginning, then many stuff improved, allowing for easier deploy, easier logging, and finally easier auto-scaling. When the ECR (AWS managed registry) was added to our region, it was quite a party @Sqreen :) I would see no point leaving it for something else today. A remaining issue is that you cannot spawn two containers speaking to a given ELB (AWS load balancer) on the same host if they need to bind the same port.
- tjholowaychuk 10y agoI do think they need to put more effort on CLIs etc, instead of relying on OSS to fulfill this niche, or at very least put more effort into supporting OSS. Lambda is similar, we have 'Serverless' and I'm hacking on Apex (https://github.com/apex/apex https://github.com/apex/apex) just to make it usable. I get that they want to create building blocks, but at the end of the day consumers just want something that works, you can still have building blocks AND provide real workable solutions. I was part of the team migrating Segment's infra to ECS, and for us at least it went pretty well, some issues with agents disconnecting etc I sort of wrote off since ECS was so new at the time. Another annoying thing not mentioned in the article is that the default AMI used for ECS is not at all production ready, you really have to bake your own images if you want something usable. I suppose this is maybe because there's subjectively no "good" defaults, I'm not sure, but it's a bit of a pain. ELB for service discovery is fine if you can afford it, I had no issues with that, ELB + DNS keeps things very simple. I'm not a huge fan of all these complex discovery mechanisms, in most cases I think they're completely unnecessary unless you're just looking to complicate your infrastructure. I also think in many cases not propagating global config (env) changes, is a good thing, depending on your taste. Scoping to the container gives you nice isolation and and more flexibility if you need to redirect a few to a new database for example. You don't have to ask your-self "shit, which containers use this?", it's much like using flags in the CLI, if we _all_ used environment variables in place of every flag it would be a complete mess. EDIT: I forgot to mention that the ELB zero-downtime stuff was awesome, if you try and re-invent that with haproxy etc, then... that's unfortunate haha. No one should have to implement such a critical thing.
- maslam 10y ago>Another annoying thing not mentioned in the article is that the default AMI used for ECS is not at all production ready, you really have to bake your own images if you want something usable. I suppose this is maybe because there's subjectively no "good" defaults, I'm not sure, but it's a bit of a pain. We ran into this as well - I forgot to add this to the post. The Amazon Linux AMI for ECS has _very specific defaults_ that need tweaking.
- nathanboktae 10y ago> I also think in many cases not propagating global config (env) changes. ... You don't have to ask your-self "shit, which containers use this?", I agree and we (dev lead at Appuri here) achieve the best of both worlds from Kube by in the secrets section of a deployment definition, specifying what secrets we need, but not the value. So we know what services need it, and it's updated in one place. That's just for the secret store though, but we could put non-secrets in secret to use that mechanism.
- dperfect 10y ago> ECS doesn't have a way to pass configuration to services I believe this is the recommended way: ECS container instances automatically get assigned an IAM role[1], with credentials accessible via instance metadata (169.254.169.254) [2]. Containers can access that metadata too. The AWS SDK automatically checks that metadata and configures itself with those credentials, so all you have to do is give your IAM role access to a private S3 bucket with configuration data and load that configuration when booting up your app. That way there's no need to copy/paste variables, and no leaking secrets in ENV variables. You do have to be careful though (as with any EC2 instance) not to allow outside access to that instance metadata endpoint, e.g., in a service that proxies requests to user-defined hosts on the network (but if you're doing that, you've got a lot more to worry about anyway). [1] http://docs.aws.amazon.com/AmazonECS/latest/developerguide/instance_IAM_role.html http://docs.aws.amazon.com/AmazonECS/latest/developerguide/i... [2] http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/iam-roles-for-amazon-ec2.html http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/iam-roles...
- embiggen 10y agoOne reason I am hesitant to go this route is because I don't want to hard-code Amazon's API's into my apps..
- deleted 10y ago[deleted]
- dperfect 10y agoI understand the reluctance to add extra dependencies (especially environment-specific ones), but in the case of a typical Ruby app, it amounts to the 'aws-sdk' gem and 1 or 2 lines in an initializer. For my own purposes, I weighed that against the alternatives[1], and it seems like a fairly reasonable compromise[2]. That won't be the case for everyone, obviously. [1] http://elasticcompute.io/2016/01/21/runtime-secrets-with-docker-containers/ http://elasticcompute.io/2016/01/21/runtime-secrets-with-doc... [2] I'm referring specifically to passing secrets (or other static values) into a container, since that seems to be what the author was talking about. For configuration requiring more complexity, of course other tools are probably more appropriate. In that case, it's outside the scope of what I would reasonably expect ECS to do.
- cxmcc 10y agoOur experience with ECS (at instacart) is not the best but we managed to get it work. Here is how we get around the issues mentioned in the article: * Service discovery: built our own with rabbitmq (we use that before ECS anyway). * Configs: pass a s3 tarball url as environment variable, download it in containers. * Cli: built our own with help of cloudformation * Agent disconnecting: we did not see situation where all agents disconnected. we use a large pool of instances, there was never an issue to start containers because of agents. In addition to these, we also do the following to make ECS work as we want it to: * built our blue-green deploy solution (structure provided by ECS is very limited) * built our own solution to integrate with ELB (ELB allows only one port per ELB)
- graffitici 10y agoAnybody has insights about using Docker Swarm? I imagine Kubernetes has been battle-tested way more in production, especially by the likes of Google. But from what I understand, Docker is really pushing swarm. I'd be curious to hear if others even considered Swarm before choosing K8s..
- lobster_johnson 10y agoThere's not really any comparison. Docker is clearly beefing up Docker/Swarm to be more like Kubernetes, but in its current state, Swarm is just a glorified Docker Compose. For example, it does not handle services (K8s can automatically provision a load balancer against all your containers), there's no volume handling, no centralized logging, no label-based targeting, it has very limited scheduling (K8s uses cAdvisor to help scheduling, can automatically ensure that pods are spread out across multiple AZs, etc.), etc. It'll be interesting to see what happens as Docker starts pushing into Kubernetes' space. Given the multiple points of overlap/contention between K8s and Docker (you have to disable Docker's built-in networking and iptables management; Kubelet has to continually monitor Docker for orphaned containers and volumes and so on; etc.) I wouldn't be surprised if Google one day decides to eliminate the Docker daemon as a dependency entirely, by writing a bare-bones container engine into Kubelet.
- smarterclayton 10y agoNot really Google driving this, but there is active work on integrating the OCI runtime (the "standard" evolution of libcontainer from Docker, used in Docker 1.12) as a container runtime to Kube. The desire is to reduce some of the overlap between container daemons on the nodes, but also support a wider array of container runtime setups (being able to run VMs via hyper, rkt containers, OCI containers, Docker containers, etc). Each of those mentioned technologies is being sponsored by different parts of the Kubernetes community, but the goal is to have more power and flexibility at runtime. Docker will continue to be a primary part of the story.
- nakagi 10y agoReally? I also think Docker Swarm Mode is still behind of K8S, but as far as I read the doc, they support - load balancing between container - volume handling https://docs.docker.com/engine/tutorials/dockervolumes/ https://docs.docker.com/engine/tutorials/dockervolumes/ - label-based constraint I know some features are not so sophisticated compared with K8S and there is no AZ awareness, but Swarm may try to catch up with it.
- nzoschke 10y agoThanks for the shoutout to Convox! I'm on the core team. I understand these challenges. I wrote about a lot of them here: https://convox.com/blog/ecs-challenges/ https://convox.com/blog/ecs-challenges/ But we have been having tons of success on ECS both for our own stuff and for hundreds of users. I see the agent disconnection problem too. convox automatically marks those as unhealthy and the ASG replaces them. It's happening more than I'd like but I'm seeing little to no service disruption. One of the root causes is the docker daemon hanging. Glad Kubernetes is working well for you. Many roads lead to success as the cloud matures.
- maslam 10y agoThat's a great blog post. Thanks for sharing!
- siliconc0w 10y agoWe evaluated ECS and Beanstalk but ended up writing a tool around building CoreOS/Fleet clusters (not currently opensource but I'm trying). We ran into similar complaints. CoreOS comes with Etcd which though initially unstable is now solid and incredibly handy for service discovery and configuration. We're using https://github.com/MonsantoCo/etcd-aws-cluster https://github.com/MonsantoCo/etcd-aws-cluster to configure it dynamically. We use etcd+confd to drive nginx containers for routing. All in all it works well. Our biggest problems are docker bug related and those we can generally handle by just terminating the node and letting autoscale heal the cluster.
- pbkhrv 10y agoWe switched from ECS to Docker Cloud and never looked back.
- rjurney 10y agoRunning DCOS (data center operating system) on AWS is a snap, and solves all these problems. It makes running docker images a no-brainer compared to all other solutions, and this includes docker images that interact with one another (not just 100 apache servers or whatever). It is the best software I have ever used, hands down. It is the second coming of zeus buddha jesus belly. It makes scaling anything in the cloud easy. No, I do not work there. No, I am not exaggerating. Yes, I spent a month fucking with swarm and service discovery before deploying a large cluster of my service in two days on DCOS. Docker is stuck in the 'one image on one machine' mindset. DCOS is taking over at the higher levels of the stack. Mark my word. https://dcos.io/ https://dcos.io/
- maslam 10y ago@rjurney - we started with ECS right around when DCOS was coming out of alpha (?). Anyway, it looks slick!
- deleted 10y ago[deleted]
- Girlang 10y agoIs this more of a "Docker" problem? Maybe Docker isn't ready for prime time.
- jordyjordy 10y agoemail demonteco@outlook.com.com for your university grades hack, pay-pal hack, whatsapp, bbm, telegram and viber hack, western union hack, social media hack, phone calls hack, mails hack, credit card hack and bank account hack, IOS hack and data recovery erase criminal records, delete employer’s bad comment, we also tutor and give free ebooks on how to become a professional hacker
- x0rg 10y agoI hope the future of cloud will really be managed OSS as service. Google is doing a great job with Kubernetes and GKE and I hope the other providers will understand that. Microsoft is on the right way with DCOS as service, Amazon is just not there yet.