11 ms·
When AWS Autoscale Doesn’t
- avitzurel 8y agoThere are many limitations that you need to "read between the lines" with AWS auto scaling. For example, we have daemons reading messages from SQS, if you try to use auto scaling based on SQS metrics, you come to realize pretty quickly that CloudWatch is updated every 5 minutes. For most messages, this is simply too late. In a lot of cases, you are better off with updating CloudWatch yourself with your own interval using lambda functions (for example) and let the rest follow the path of AWS managed auto scaling. There is also a cascading auto scale that you need to follow. If we take ECS for example, you need to have auto scaling for the containers running (Tasks) AND after that you also need auto scaling for the EC2 resources. Both of these have different scaling speeds. Containers scale instantly while instances scale much slower. Even if you pack your own image, there is still a significant delay.
- avitzurel 8y agoFrom a bird's eye view, you also need to figure out what costs you more. For example (and I know nothing about the use-case of OP, can only estimate), you might be able to buffer requests into a queue and have it scale up slower. You might have auto scaling that needs to be close to real time and auto scaling that can happen on a span of minutes. Every auto scaling needs to also keep in mind the storage scaling, often you are limited by the DB write capacity or others.
- eunoia 8y agoOut of curiosity, what’s the use case for running ECS on EC2 (instead of using Fargate) these days?
- lukeck 8y agoApart from pricing and the potential to overcommit resources on EC2-ECS, there are a couple of other differences. One is your options for doing forensics on Fargate. AWS manage the underlying host so you give up the option of Doing host level investigations. It’s not necessarily worse as you can fill this gap in other ways. Logging is currently only via CloudWatch logs so if you want to get logs into something like Splunk you’ll have to run something that can pick up these logs. You’ll have that issue to solve if you want logs from some other AWS services like Lambda to go to the same place. The bigger issue for us is that you can’t add additional metadata to log events without building that into your application or getting tricky with log group names. On EC2 we’ve been using fluentd to add additional context to each log event like the instance it came from, the AZ, etc. Support for additional log drivers on Fargate is on the public roadmap[1][2] so there will hopefully be some more options soon. [1] Fargate Log driver support v1 https://github.com/aws/containers-roadmap/issues/9 https://github.com/aws/containers-roadmap/issues/9 [2] Fargate log driver support v2 https://github.com/aws/containers-roadmap/issues/10 https://github.com/aws/containers-roadmap/issues/10
- etaioinshrdlu 8y agoPerhaps cost?
- jusssi 8y agoAt least one is, that you get to use the leftover CPU and memory for your other containers when you use an EC2 instance. With some workloads this lets you to overcommit those resources if you know all your containers won't max out simultaneously. Edit: another one is that you can run ECS on spot fleet and save some money.
- matwood 8y agoA large part of our ECS capacity is running on spot instances which are much cheaper.
- NathanKP 8y agoAWS employee here. If you are able to achieve consistent greater than 50% utilization of your EC2 instances or have a high percentage of spot or reserved instances then ECS on EC2 is still cheaper than Fargate. If your workload is very large, requiring many instances this may make the economics of ECS on EC2 more attractive than using Fargate. (Almost never the case for small workloads though). Additionally, a major use case for ECS is machine learning workloads powered by GPU's and Fargate does not yet have this support. With ECS you can run p2 or p3 instances and orchestrate machine learning containers across them with even GPU reservation and GPU pinning.
- chrissnell 8y agoI'm not totally up to speed on ECS vs EKS economics but it seems like EKS with p2/p3 would be a sweet solution for this. Even better if you have a mixed workload and you want to easily target GPU-enabled instances by adding a taint to the podspec.
- NathanKP 8y agoKubernetes GPU scheduling is currently still marked as experimental: https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus/ https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus... ECS GPU scheduling is production ready, and streamlined quite a bit on the initial getting started workflow due to the fact that we provide a maintained GPU optimized AMI for ECS that already has your NVIDIA kernel drivers and Docker GPU runtime. ECS supports GPU pinning for maximum performance, as well as mixed CPU and GPU workloads in the same cluster: https://docs.aws.amazon.com/AmazonECS/latest/developerguide/ecs-gpu.html https://docs.aws.amazon.com/AmazonECS/latest/developerguide/...
- avip 8y agoCost and legacy.
- takeda 8y agoFargate is orthogonal to ECS and can be used together. The difference is that instead spinning VMs as hosts and configuring them for ECS and worrying about spinning just right amount of them, you select fargate, which does all of that behind the scenes (kind of like lambda), but the VM instances provided by fargate are a bit more expensive.
- jcrites 8y agoThe effectiveness of dynamic scaling significantly depends on what metrics you use to scale. My recommendation for that sort of system is to auto-scaled based on the percent of capacity in use. For example, imagine that each machine has 20 available threads for processing messages received from SQS. Then I'd track a metric which is the percent of threads that are in use. If I'm trying to meet a message processing SLA, then my goal is to begin auto-scaling before that in-use percentage reaches 100%, e.g., we might scale up when the average thread utilization breaches 80%. (Or if you process messages with unlimited concurrent threads, then you could use CPU utilization instead.) The benefit of this approach is that you can begin auto-scaling your system before it saturates and messages start to be delayed. Messages will only be delayed once the in-use percent reaches 100% -- as long as there are threads available (i.e., in-use < 100%), messages will be processed immediately. If you were to auto-scale on SQS metrics like queue length, then the length will stay approximately zero until the system starts falling behind, and then it's too late. If you scale on queue size then you can't preemptively scale when load is increasing. By monitoring and scaling on thread capacity, you can track your effective utilization as it climbs from 50% to 80% to 100%; and you can begin scaling before it reaches 100%, before messages start to back up. The other benefit of this approach is that it works equally well at many different scales; a threshold like 80% thread utilization works just as well with a single host, as with a fleet of 100 hosts. By comparison, thresholds on metrics like queue length need to be adjusted as the scale and throughput of the system changes.
- xfitm3 8y agoI'm surprised I didn't see application performance monitoring mentioned here. A lot of applications are complex and in those cases adding containers is only effective until you reach the next constraint. Having two resources (such as DB and app) scale in concert can be exceedingly difficult.
- guywithabike 8y ago(Author here) Absolutely! It's amazing how complex things get on the configuration side as you try to get smarter about autoscaling. The big downside of trying to be too clever here is that you wind up with a wildly brittle autoscaling setup that falls over as soon as the underlying assumptions around the relationship between your metrics and your scaling needs change. As in many things engineering, we've found that it's best to keep the configuration as simple as possible and use a solid foundation of reporting / alerting to give you an early heads up that you need to revisit and update your autoscaling strategy.
- eikenberry 8y ago> Having two resources (such as DB and app) scale in concert can be exceedingly difficult. This means your resources are too tightly coupled. If they are so tightly coupled that they need to scale together then they are not two resources and you should look into restructuring them into two actual resources or bind them more closely to make a single resource.
- xfitm3 8y agoIn my example of DB and backend: How could they be decoupled?
- eikenberry 8y agoAs far as runtime, applications and DBs are already decoupled. You have N application instances mapped to M database instances. Applications can usually scale pretty much with load. Databases vary wildly in how they scale and it depends on the DB type.
- 8y ago
- tyingq 8y agoVertical autoscale is on my wish list. Some way to automatically scale instance size for those things that don't scale well horizontally.
- barbecue_sauce 8y agoI didn't even realize this wasn't part of the offering yet.
- deleted 8y ago[deleted]
- xorgar831 8y agoJoyent offers this as I recall, however it's only a scale up, I don't think they can scale down afterwards.
- jedberg 8y agoYou could hack around this by creating two auto-scaling groups with different instance types and then have them follow the same metric such that the small group goes to 0 and the larger one spins up. Not a great solution but better than nothing.
- etaioinshrdlu 8y agoJelastic claims to do this, and the marketing made it sound so cool. https://jelastic.com/ https://jelastic.com/ But when I tried it, it turns out the docker support requires very specific base images. Not really docker then is it?!?
- doctorpangloss 8y agoAn incredible amount of software and infrastructure is written precisely for analytics data gathering workloads. I'm pretty confident AWS's product for this use case would be Lambda and the new on-demand DynamoDB. Is there actually a use case in analytics that requires a server that accepts connections from multiple clients, and then has to have <60ms latency including state over the wire and executing sophisticated business logic, between those clients, for time periods longer than 5 seconds? I.e. something that resembles a video game? Because if there isn't, if your goal is to scale, why have containers at all?
- thecopy 8y agoBatching incoming requests, for one. Kinesis only allows 5 write requests per second per shard, for example. As well, Lambda have limits regarding concurrent executions and are very slow (10s) if needing VPC connectivity (in this case the default concurrent lambda limit is 350 due to ENIs)
- doctorpangloss 8y agoI suppose you could just ask them to raise the limits?
- deleted 8y ago[deleted]
- meekins 8y agoHmm... I don't see anything in the docs implying that - Kinesis API docs say it's possible to ingest 1000 records or 1MB per shard per second. There's a 5/s limit on reads however but those deal with batches of records anyway. We have one service running that consumes data to a Kinesis stream published as an API GW endpoint. Preprocessing is done in Lambda in batches of 100 records and the processed records get pushed to Firehose streams for batched loads to a Redshift cluster for analytics. So far we've been very happy with the solution - very little custom code, almost no ops required and it performs and scales well.
- Legogris 8y agoHaving used ECS quite a bit, I do not recommend anyone building a new stack based on it. Kubernetes solves everything ECS solves, but usually better and without sveral of the issues mentioned here. Last time I checked, AWS was still lagging behind Azure and GCP on Kubernetes, but I have a strong feeling they're prioritizing improving EKS over ECS. If you're already invested in ECS it's a different story, of course.
- pram 8y agoI think that Fargate is the “improvement” for ECS. I never understood the appeal of ECS in the first place, seemed (and still does) really half baked.
- taormina 8y agoFargate is almost what the marketing team said ECS was going to be.
- NathanKP 8y agoAWS employee here. Sorry to hear that you feel ECS is half baked. Feel free to reach out directly using the details in my profile info if you have any feedback you'd like me pass on to the team. To clear up the confusion on the relationship between Fargate and ECS, think of Fargate as the hosting layer: it runs your container for you on demand and bills you for the amount of CPU and GB your container reserved per second. On the other hand ECS is the management layer. It provides the API that you use to orchestrate launching X containers, spreading them across availability zones, and hooking them up to other resources automatically (like load balancers, service discovery, etc). Currently you can use ECS without using Fargate, by providing your own pool of EC2 instances to host the containers on. However, you can not use Fargate without ECS, as the hosting layer doesn't know how to run your full application stack without being instructed to by the ECS management layer.
- pram 8y agoFrom my perspective Fargate offers the functionality I would have expected from ECS in the first place. What ECS provides OOB requires too much janitoring and ultimately isn't terribly different in effort compared to running your own k8s or mesos infra on EC2 instances you provisioned yourself. You still basically needed an orchestration layer over ECS. Which is why, I assume, Fargate is now listed as an integral feature of ECS on the product page.
- b5u 8y agoI've also been deploying services on ECS for close to a year now and would like to address some inaccuracies the author seems to have made: 1) in 'Surprise 1' the author offers examples of CPU Utilization (or target) is between 80% and 95% without mentioning the reserved CPU/memory (aka size) of those tasks (under the assumption that he's using the Fargate launch type). The 'size' of a task also influences the average CPU target utilization. For instance, if a task requires the reserved CPU of 4 vCPUs, then a spike from 80% to 95% is handled differently than when a task reserves 1 or 2 vCPUs. The same goes for memory. In an example setup I'd use 1-2 vCPUs sized tasks with a service-wide target avg. CPU Utilization of 70% along and a StepScaling policy which adds 10% more tasks if the service avg. CPUU falls between 70-80, 20% if between 80-90 and 25% if above 90. My strategy has been being smaller-sized tasks, lower service avg CPU utilization (compared to 80%-90%) and shorter evaluation periods/datapoints for the scale-out CW alarms (minimum being 60 seconds IIRC). The short evaluation periods/low number of datapoints of the CW alarm allowed me to handle spikes reasonably fast. 2) in 'Surprise 3' the author claims that the Terraform's aws_appautoscaling_policy 'is rather light on documentation'. Since I am a user of Terraform for several years, I find it inaccurate mostly because of the several examples available in the documentation https://www.terraform.io/docs/providers/aws/r/appautoscaling_policy.html https://www.terraform.io/docs/providers/aws/r/appautoscaling... as well as many more when doing a Github exact search for "aws_appautoscaling_policy" language:HCL will reveal many, many more examples from open-source repos (some with permissive licenses too). I'd created a custom ecs-service TF module which creates for each service (optionally) an ALB along with listeners and the attached ACM-issued TLS certs and TGs, the scale-in/out CW alerts with configurable thresholds/policies, SGs, Route53, etc. allowing one to quickly configure and launch an ECS service fast and reliably. Regarding the scale-in, I typically also have that at intervals between 5-15 minutes to avoid an erratic scale-in/scale-out 'zig-zag' happening even at the cost of briefly over provisioning.
- argd678 8y agoThe biggest scaling issue I always run into is the database is the bottleneck and there’s not a lot of options for most databases to auto scale them.
- brianwawok 8y agoYup. Your DB usually has to be over-provisioned for peak WRITE capacity. Read capacity is easy to skill to infinity with caches. But if a DB can only write 1000 updates per second, nothing will change that. In many cases - it's ok to not process EVERYTHING right away. Process the important stuff RIGHT AWAY. Slowly process the unimportant stuff in your spare time.
- DVassallo 8y agoThe way I've been happiest using EC2 Auto Scaling was to have a single cron-job continuously calculating how many instances I should be running, and it sets the desired capacity manually with the Auto Scaling API[1]. This may seem to defeat the purpose of Auto Scaling, but it's actually much more convenient than spinning up/down EC2 instances with the EC2 API. You get to precisely control how to scale, and won't be at the mercy of the Auto Scaling heuristics. [1] https://docs.aws.amazon.com/autoscaling/ec2/userguide/as-manual-scaling.html https://docs.aws.amazon.com/autoscaling/ec2/userguide/as-man...
- d3sandoval 8y agoWe did the same thing! Glad to hear we weren't too crazy in doing so
- teej 8y agoI don’t think it’s crazy to think you are more capable of predicting the unique demand curves of your business better than a heuristic designed to be good enough for the median Amazon customer.
- samstave 8y agoSo we had this cause a spectacular outage a few years ago. We were doing exactly this - but we had a flaw: we didnt handle the case when the AWS API was actually down. So we were constantly monitoring for how many running instances we had - but when the API went down, just as we were ramping up for our peak traffic - the system thought that none were running because the API was down - so it just kept continually launching instances. The increased scale of instances pummeled the control plane with thousands of instances all trying to come online and pull down their needed data to get operational -- which them killed our DBs, pipeline etc... We had to reboot our entire production environment at peak service time...
- robrtsql 8y agoI don't understand--how were you launching instances if the API(s) was/were down? Your system was unable to determine that there were instances running, but it was able to send RunInstances requests to EC2?
- acd 8y agoAWS needs to glue EC2 and ECS scheduling together. Today the schedulers are separate. So basically the feet does not know what the arms are doing. That leaves fixing this scaling up to the client meaning duplicate code effort solving the same thing for each AWS customer.
- avitzurel 8y agoYour last sentence is something I have been thinking about for a very long time. I have been in Devops since it before it had a name and I see many companies solving the same problems. Like ths auto-scaling post. That's not the first company to deal with it (nor the last). Providing a set of tools can be very beneficial but so hard to dial down. I have a very big itch around solving this problem.
- NathanKP 8y agoAWS employee here. This is a feature that is currently on our public roadmap for container services, in the "Researching" category: https://github.com/aws/containers-roadmap/issues/76 https://github.com/aws/containers-roadmap/issues/76 Feel free to drop a thumbs up on the roadmap item to show your support and boost its priority on the roadmap, or leave a comment to let us know more about your needs.
- jedberg 8y agoAuto-scaling is depending on startup time. If your startup time for a new instance/container is 5 seconds, then you need to predict what your traffic will be in 5 seconds. If your startup time is 10 minutes, then you need to predict your traffic in 10 minutes. The choice of metric is important, but it needs to be a metric that predicts future traffic if you want to autoscale user facing services. CPU load is not that metric. The best way to do autoscaling is to build a system that is unique to your business to predict your traffic, and then use AWS's autoscaling as your backup for when you get your prediction wrong.
- k__ 8y ago°whispers° cloud native...!
- ravedave5 8y agoI got bit by this! Even worse is that one of the servers crumpled because we didn't scale up fast enough - so AWS killed it because of the health metric. Which then took out the remaining two because they were then far, far over capacity. I got the pager duty alert and found a total cluster and just manually set it to scale up way bigger. Now for all big events we manually bump minimum server counts for that period :\
- eridius 8y ago> For example, the maximum value for CPU utilization that you can have regardless of load is 100%. I'm surprised it can't use load average to estimate the true resource demand.
- wahnfrieden 8y agoYou can, but it’s not visible to the hypervisor (it’s an OS concept) so you have to publish that metric from an agent on the machine. Then you can use it for autoscaling.
- viraptor 8y agoBut even then, "load" doesn't work well for all workloads. http://www.brendangregg.com/blog/2017-08-08/linux-load-averages.html http://www.brendangregg.com/blog/2017-08-08/linux-load-avera... The number of switching tasks may be very high for a number of reasons, including a very large number of threads which do very small chunk of work each and yield.
- jugg1es 8y agoI think most people want to write this kind of blog after slogging through the AWS learning curve. Then they figure out how to use it and the urge goes away.
- auslander 8y ago> .. the ECS dashboard does not yet support .. Terraform .. They haven't yet matured enough, it seems. Cloudformation is the right way to code your infra, not web Console or anything else. Good sign, though, is they use ECS, not Kubernetes :)
- anbotero 8y agoQuick question for the AWS employee solving inquiries: I used ECS in 2017, and back then there was this weird issue where sometimes tasks would switch to new versions in like, a minute (if that), but sometimes, like 2/10, it would take like 10-12 minutes just for it to start killing old Task containers. Back then there wasn't any timeout option or anything to force the killing. Do you know if now there is? The project was killed for different reasons, but I really liked everything else on ECS. Thanks! EDIT: I meant killing containers, not the Tasks themselves. Sorry.
- Buge 8y ago>So, if you’re targeting 95% CPU utilization in a web service, the maximum amount that the service scales out after each cooldown period is 11%: 100 / 90 = 1.1 How many errors are in that sentence? The 95% -> 90 mysterious conversion. 100/90 is actually 1.111(repeating of course), not 1.1. And if it did equal 1.1 it would be 10%, not 11%.
- mnutt 8y agoThe biggest challenges I’ve had with auto scaling have been slow scaling time and default metrics not being a good proxy for scaling needs. One thing I was mildly curious about: if you’re going to build your own metrics and scaler, what would be some of the downsides of having it scale down by just putting instances in the Stopped state, then scale up by starting them? In my experience starting takes seconds while launching new instances takes minutes. Having to deploy updates to stopped instances would be complicated and you’d have to pay EBS costs for stopped instances, but I’m curious if there are other issues. Launching an instance from an AMI, even after the instance comes up the disk tends to be very slow for some time as if it’s lazily loading the filesystem over the network.
- svsucculents 8y agoThat's why you don't use AWS 'auto' scaling. Every application has it's sweet spot and it's simpler to roll your own when you know best.