7 ms·
While the new C5 instances are certainly welcome - I've been hoping for their release since their announcement in November 2016 (and they were already late for
by STRML 9y ago
While the new C5 instances are certainly welcome - I've been hoping for their release since their announcement in November 2016 (and they were already late for Skylake at that point) - we have encountered a number of show-stopping problems that point to this project being just a bit too ambitious.
To name a few:
1. EBS volumes attached to C5 instances show completely bogus CloudWatch metrics, over an order of magnitude higher than reality (e.g. average read/write latency prints at 100-60,000ms depending on load)
2. C5 instances don't work - at all - behind an NLB with a Target Group pointing to it as an "instance". You have to put it in "IP" mode.
3. OpsWorks, as always, lags way behind AWS offerings. You can't launch C5 instances. This was true of even R4 instances for a while, but you could at least change them via API. Not so with the C5 instances; unless you want to lose track of their type completely, you just have to abstain for now.
3a. As a result, we have to run R4 instances for some of our web tier - despite not needing the memory - because they have the highest network allocation. To make matters worse, AWS won't tell you the network allocation. You don't know until you start dropping packets.
4. ZFS on C5 instances can behave strangely. We've been unable to resize drives (zpool online -e <pool> <drive>) if they're identified by ID ("unable to read disk capacity"). Moving the instance back to any other type fixes the issue.
As always, you expect a couple of quirks with a new architecture, but I found myself wishing they had just stuck a new board with a Skylake chip into a rack and launched it.
Compared to this, GCE has a far more attractive offering: even ignoring all these issues, we simply can't get the instance size we want (a few fast CPUs + lots of memory). It just doesn't exist.
- tty7 9y agoYou have given me a whole other perspective to an aws customer. I spend a lot on aws (70k/m +) but id be happy if i had core2duo cpus! What sort of work do you have that requires the latest generation ? Or more why do you want the latest ? Id expect aws to always be behind the ball - are they the right platform for you? Gce is interesting, im moving half of my infra over there - but again its not really about their hardware offerings. Its more about being multi cloud/redundant
- STRML 9y agoOur particular use case involves a few processes that are heavily serial, while yet memory-intensive; the fastest possible processor would be a boon to us (for instance, if someone would guarantee an overclocked Xeon, we'd take it in a heartbeat for almost any price). AWS is indeed behind the curve more often than not, but they also have some amazing hosted products. That was a bigger deal in 2014 than it is now (Kubernetes will eat the world) but it's still a reliable offering, and reliability is our most important criteria. Agreed re: being multi-cloud in the end. It's the only responsible choice above a certain scale.
- mbell 9y ago> I've been hoping for their release since their announcement in November 2016 (and they were already late for Skylake at that point) I don't know what you mean by 'late'. Skylake Xeons were delayed and have only been released recently, at least partially. You may be able to get them from vendors like Dell but you still can't go out and buy one anywhere that I'm aware of.
- toast0 9y ago> 3a. As a result, we have to run R4 instances for some of our web tier - despite not needing the memory - because they have the highest network allocation. To make matters worse, AWS won't tell you the network allocation. You don't know until you start dropping packets. Ugh, I hate that. They also have hidden limits on the number of incoming tcp connections you can have.
- STRML 9y agoWorse still, we spent weeks on the phone with AWS insisting there was no throttling - just to find out there was throttling. And that would have been discoverable if a single engineer had looked at CloudWatch.
- otterley 9y ago> They also have hidden limits on the number of incoming tcp connections you can have. Do you have proof of this?
- toast0 9y agoNo, I had an experiment I was tried to run on EC2 in 2013, and ran into this. It was very clear though. Established connections would plateau, and then no more tcp syns would arrive to the ec2 host, unless a connection was closed. At the time the limit was 5000 connections on the micro instance -- from my notes, we got an allocation for hi1.4xlarge and I know we hit the limit there too, but I don't recall what the limit was. This was a very simple TCP proxy, using Linux kernel ipmasq; this uses significantly less memory than HAProxy, although with a lot less features. We had a very excited account rep because of where I work, but he was only barely able to confirm the limits were there, he wasn't able to get them raised or removed, nor could he tell us the limits by machine type. There's a thread from 18 months ago on the aws forums [1], where an aws rep more or less confirms, but again provides no information [1] https://forums.aws.amazon.com/thread.jspa?threadID=231806 https://forums.aws.amazon.com/thread.jspa?threadID=231806
- dastbe 9y agoSo I wouldn't say C5s don't work at all with NLB instance target groups - I'm running one right now. I wouldn't be surprised if you're hitting an edge case with the tighter integration* between NLB's and instance associations. If you haven't already, please do reach out to support. * from the docs "If you specify targets using an instance ID, the source IP addresses of the clients are preserved and provided to your applications. If you specify targets by IP address, the source IP addresses are the private IP addresses of the load balancer nodes."
- STRML 9y agoIt appears the specific issue is with an NLB in one subnet referencing a C5 in another subnet. It's useful to put the targets in a private subnet and the NLB in a public subnet, so this breaks usage for us.
- joemag 9y ago#2 is expected to work. Can you reach out to me directly (joemag@) and I will make sure we take a look.
- STRML 9y agoThere's an open support ticket, but I've reached out as well.
- _msw_ 9y agoDoes the updated documentation at http://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/ebs-metricscollected.html http://docs.aws.amazon.com/AmazonCloudWatch/latest/monitorin... explain the CloudWatch metrics?
- STRML 9y agoThat's more helpful. That's a lot of special cases - and thought - to put into each time we look at a C5-attached volume as opposed to any other. It also does not include any mention of Average Read Latency and Average Write Latency, which appears to be incorrect in all dimensions (average, min, max, sum).
- otterley 9y agoWe've found that CloudWatch is useful for some system and I/O related metrics (i.e., the from the EBS SAN's point of view) but less useful for others (i.e., metrics from the VM's point of view). It's worthwhile to install a monitoring agent on the VM that can collect I/O latency and other statistics there too. We use Datadog, but there are lots of options out there (collectd etc.).
- _msw_ 9y agoThis is a topic that we're continuing to iterate on between EBS, CloudWatch, and the AWS console team for metrics. When a newly introduced behavior makes step function changes in graphs displayed in the AWS console it doesn't meet the principle of least astonishment. We're also continuing to investigate the reported latency in the console. Because this is a derived metric from VolumeTotal{Read,Write}Time and Volume{Read,Write}Ops, there may be a miscalculation happening due to the change in dimensions.
- cperciva 9y agoNot sure if this is related, but when I was getting EFS working last year I noticed that CloudWatch graphs were often complete nonsense due to EFS not logging zeroes for idle filesystems but CloudWatch treating this as "missing data" rather than "implied zeroes".