6 ms·
Tell HN: AWS connectivity issues, but health dashboard says everything fine
About 15 minutes ago I got a call from a customer that our site was down. "Works for me," I said, because I could bring it up on my laptop. The customer said he couldn't get it on his phone, and then I confirmed I couldn't get it on my phone either.
The AWS Health Dashboard (https://health.aws.amazon.com/health/status) reports no issues at all. But DownDetector (https://downdetector.com/status/aws-amazon-web-services/) shows a spike in reports.
I can't even reach the AWS console through my phone.
So, AWS has connectivity issues to certain networks and their own health dashboard is lying to us about it. What gives?
(All of this accurate as of 2:10 pm CST).
Update: as of 2:26 pm CST, the health dashboard reports that they are "investigating an issue". So, 45 minutes after Down Detector sees it, they do.
- xup 4y agoI'm still seeing it on my end. Our currently-running EC2 instances are working fine, but the EC2 us-east-2 console webpage doesn't load, and an EC2 instance in us-east-2 I rebooted has yet to come back online.
- the_snooze 4y agoI'm seeing it too, and surprised their health page says nothing. The US East 2 console is unresponsive. https://us-east-2.signin.aws.amazon.com/ https://us-east-2.signin.aws.amazon.com/
- mrobins 4y agoWe're experiencing issues for some users but not all connecting to resources in US East 2.
- muttantt 4y agoWe got all our main stack all at us-east-2. Seems to all be running currently
- ccleve 4y agoTry it on your phone.
- joshuanapoli 4y agoWe couldn't detect any problem accessing resources via CloudFront/AppSync. Maybe the issue was specific to ELB.
- rshm 4y agoCould be resolved, east-2 console working fine on my end.
- jamroom 4y agoWe're still up on us-east-2 but lots of customers calling in that they can't connect - makes me think there's some network down somewhere.
- cathintexas 4y agoOn our team, we are seeing that if you are on AT&T cell service or have AT&T as your ISP, you can't reach AWS or our site in US-east-2.
- mplanchard 4y agoWe're seeing a similar thing for our us-east-2 properties. Some of our team is able to reach them, but others aren't. Folks in the midwest (Oklahoma and Michigan) can't even load the AWS console, while people in Texas, California, Arizona, and Pennsylvania can.
- nnf 4y agoI'm hearing from customers and other employees that our stuff at AWS (us-east-2) is unreachable, but I'm able to get to it all without any issue (via http & ssh). Perhaps there's a problem upstream of AWS that's only affecting some ISPs?
- thedougd 4y agoIt's only impacting some ISPs. Outage varies by my office locations. At my location, it is out, however I was able to get access via Cloudflare warp VPN. Edit: Sounds like AT&T
- deleted 4y ago[deleted]
- BWStearns 4y agoSeeing issues in Florida for us-east-2. Coworkers in NY can still get to us-east-2.
- _justinfunk 4y agoIt does seem to be a networking issue. I have a ec2 instance in us-east-2 that is accessible through a "Global Accelerator" but not externally through my ISP. That ec2 instance can talk to other ec2 instances that are on us-east-2 - but none of those other instances are accessible externally.
- leesalminen 4y agoCan confirm that Global Accelerator helped us avoid this issue today.
- mrobins 4y agoConfirmed by AWS: https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status
- jstimps 4y agoWe're seeing that external requests to ALBs on us-east-2 are affected
- WFHRenaissance 4y agoCyber attack?
- jeremib 4y agoThey've finally updated their status. 12:26 PM PST We are investigating an issue, which may be impacting Internet connectivity between some customer networks and the US-EAST-2 Region.
- biggerChris 4y ago
- AlphaWeaver 4y agoAppears to affect ELBs.
- dixie_land 4y agoKeep in mind AWS status dashboard solely reflects the product owning managers discretion. And the number of yellow ("green I" if you're old enough) is definitely a material input to PIP :)
- joecot 4y agoCorrect. No matter how down AWS is, their status page will only show a disruption if a manager approves showing a disruption. There is nothing automated to display the status, so the status page is mostly worthless except for whatever AWS admits is down.
- whoknew1122 4y agoAll this compounded by the fact AWS builds on AWS. So there can be a disruption of a service, but it's not really the service's fault -- it's a upstream failure.
- hedgehog_irl 4y agoEveryone seems to overlook the point here. That yet again Amazon were slow as hell to be honest with their customers. I get it up down reports help but why do you keep using a service which lies to you about availability. I've read on HN in the past how the dashboard can only be updated to reflect an issue with approval. (Comments section on a similar posting, believe it if you wish). So why not move to a hosting company that is transparent and open about their status. I'll not make suggestions as I don't want to be accused of trying to shill for a specific provider but there are plenty out there. 45 min to update their public dash is too slow. They either don't care, don't monitor or they are trying to hide their stats for fear Jeff will beat the staff for SLA violations. If any other provider lied to customers the way AWS does they wouldn't be tolerated why do you tolerate this behaviour from AWS? Edited to fix auto correct issues
- ioman 4y agoAWS is the 800 pound gorilla in the cloud space. Are any of the other cloud providers better with customer honesty?
- autotune 4y agoAlso good luck trying to convince your company to migrate to another cloud provider over, say, implementing multi-region strategy, which you should have been doing in the first place.
- hedgehog_irl 4y agoHighlight the lack of transparency on reporting outages and that's a start. If your MPLS or ISP provider operated in the save way. The company wouldn't accept it
- autotune 4y agoMy company is not going to spend hundreds of thousands of dollars or more, and months or even years of effort, and add additional constraints to the given pool of candidates we are hiring for, to migrate to GCP or Azure or DigitalOcean or Hetzner or wherever is considered more trendy than AWS right now due to "a lack of transparency" lmao. I would look completely incompetent to even suggest the idea to anyone internally.
- Analemma_ 4y agoThere are three kinds of lies: lies, damned lies, and cloud status dashboards.
- baq 4y agoIt only turns yellow if the datacenter gets flooded by lava. Red is probably a tactical nuke.
- agilob 4y agoI read here on HN before that yellow requires a manual signature on a paper from a higher manager. Because such fault affects their compensation and decreases stock value. Red requires signature from C-level. It's not automated at all, almost worthless dashboard.
- agilob 4y agoAnyone with affected RDS instances? We were getting random connectivity issues today occasionally... New pods with 1-2 minutes after startup were suddenly getting timeouts connecting to MySQL DBs
- ocdtrekkie 4y agoA vendor's cloud product is having significant issues. Figured HN would tell me which major public cloud infrastructure fell over to cause it. Never fails.
- bloaf 4y agoThere was definitely something going on last night too, I noticed a number of sites having intermittent issues confirmed by down detector.
- daneel_w 4y agoAbout two weeks ago all three of our Aurora DB instances in eu-central-1 suddenly crashed and were offline, to no avail, for almost 55 minutes. Simultaneously we had random network problems going on within our eu-central-1 VPC which we were unable to diagnose. We still don't know what happened because we're not getting any answers to our support request. The AWS health dashboard was all green the entire time. No notifications were sent out.
- rshm 4y agoSnowflake confirmed AWS us-east-2 issues as well. AWS - US East (Ohio): INC0073093 https://status.snowflake.com/incidents/yv40l966krl9 https://status.snowflake.com/incidents/yv40l966krl9
- deleted 4y ago[deleted]
- everfrustrated 4y agoA reminder that the public and personal health dashboards are not the the only port of call. If you pay for the top tier of AWS support, if you have a suspected outage you'd be paging in AWS who will pick up the phone and start debugging your problem. If your business depends on AWS you don't sit around clicking refresh on a status page hoping it might be updated.
- coredog64 4y agoAt some level of spend, your account team will know what services you use and know when those services are having higher than normal error rates.