4 ms·
My $0.02... #1 Make EC2 Auto Scaling decisions faster - it's painfully out of sync with the AWS Console table. I don't mind paying for it. And while we're at i
by qxmat 5y ago
My $0.02...
#1 Make EC2 Auto Scaling decisions faster - it's painfully out of sync with the AWS Console table. I don't mind paying for it. And while we're at it please can I get a metadata address REST API to signal CONTINUE?
#2 SSH keys are configured in cloud-init, there's no reason they can't be read from SecretsManger and rotated out of the box.
#3 Put $HOSTNAME/$env:ComputerName in the EC2 Console table - weird software like Azure AD Application Proxy uses it for registration and I'm fed up using SSM to query all Windows instances to find an errant machine.
... and one more. Please make CloudWatch an out-of-the-box sink for cloud-init logs. SSM Agent doesn't start on Windows until fairly late (after other services), so it can be pain reading logs out of %TEMP% if you've locked the network down and fully adopted SSM sessions/port-forwarding.
... and a final one - we need a "1-click WTF" button to help solve EBS/KMS issues. There's insufficient info in the UI. Dropping down to check the cli JSON output isn't a user-friendly solution. Azure does this better with their "Diagnose and solve problems" blade in App Service.
- mdaniel 5y agoI was going to comment about whether you were aware of SSM after reading the ssh keys one, only to be surprised to see the rest of them reference SSM We don't even add a break-glass KeyName since SSM is available on all client OSes we care about, Ansible has it as a "connection:" mechanism if we need to do things to a lot of instances, and kubelet for everything else I hear you about the cloud-init output, though I would love if they used some browser trickery to grab control-w while an SSM session was active, though, because muscle memory is a harsh mistress :-(
- raffraffraff 5y agocontrol-w is my least favourite twitch. I use it randomly, often. It's not too bad in a Google Doc because you just control-shift-t.
- raffraffraff 5y ago> #1 Make EC2 Auto Scaling decisions faster Fuck yeah. The autoscaler is almost certainly a distributed, fault tolerant, highly available system that is deployed across multiple availability zones. Presumably they trade consistency for other attributes so there's an intentional delay to avoid making constant flapping changes. There's also the metrics granularity to consider. Take SQS for example, another distributed system. Many of the useful queue metrics are averaged over time, which adds a delay on top of the metric gathering period. The result is that a scale-to-zero ASG that scales based on SQS queue depth takes about 3 minutes to produce an EC2 instance from a cold state. If you're lucky.
- WJW 5y agoWhile all the things you mentioned are correct, I would also like to point out one of the fundamental results from control theory here: The time period of the control loop must be proportionate to the time it takes for any changes to the controlled system to propagate, otherwise you will get wild oscillations (in the number of instances in this case). Since it can take a few minutes for an EC2 instance to fully boot up, start up all the applications it needs and start serving traffic, you can't have a very fast autoscaler loop adding a new instance every second: it would start up way more instances than needed and eventually have to shut most of them down again, etc. This EC2 startup delay is the core problem, if you could reduce that interval the autoscaler could run way faster even with the usual distributed HA, CAP theorem and metrics concerns.
- qxmat 5y agoI suppose what I really need is a way of invoking certain operations through the ASG interface immediately, not at end of a long distributed async chain.
- mdaniel 5y agoI haven't personally tried it, but I wonder if that's what "Warm Pools" is partially designed to help? https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-warm-pools.html https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-au... I guess whether that helps the SQS situation depends on whether it is _the ASG_ decision time that's the problem, or the ASG requests an instance which itself takes multiple minutes to come into service
- hnjst 5y ago> #2 SSH keys are configured in cloud-init, there's no reason they can't be read from SecretsManger and rotated out of the box. If you didn't know it, you may be interested by ec2-instance-connect (https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-instance-connect-methods.html https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-inst...). It's sadly officially only supported on Amazon Linux / Ubuntu but ephemeral ssh key authorization based on IAM has nice properties in terms of security / auditability / access control / revocation etc.
- mdaniel 5y agoI would have thought them constraining it to Amazon Linux and (bizarrely) Ubuntu meant there was something inherent to those cloud AMIs which enabled such a thing, but this[0] sounds like just as much work as installing the SSM agent, and having the extra drag of needing to monkey with sshd_config (a fine way to lock oneself out if not careful) 0 = https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-instance-connect-set-up.html https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-inst...