17 ms·
Amazon EC2 and RDS in US-EAST zone down
multi-az deployments affected too.
- cupcake_death 14y agoYep - Forums are exploding
- mattwdelong 14y agoIt's not entirely down as I can still access my instances. I'm in us-east-1b.
- grourk 14y agoYour us-east-1b might be my us-east-1a.
- dwhsix 14y agoKeep in mind AZs are different per account. My us-east-1b is not necc'ly your us-east-1b (as someone reminded me on twitter just now).
- ahmedaly 14y agoMy instances are not down too.. I will back it up now in case things go bad.
- ahmedaly 14y agodotcloud was down also but its now up. (they rely on ec2)
- NathanKP 14y agoI am experiencing two out of four instances in us-east-1e unreachable.
- misiti3780 14y agomy instances in us-east-1c are fine
- malachismith 14y agoGoat rodeo.
- bad_user 14y agoI got notified by Pingdom that my domain was down before AWS had any info on that status page of theirs. IMHO, they should improve on the latency of their alerts.
- iharris 14y agoSNS sent me an e-mail of my instance alarms pretty quickly. EDIT: My status checks were slow to update like the sibling comment stated, although the alarms that measure system resources triggered almost immediately when everything blew up. I think the status checks refresh at a certain interval, but those aren't really meant for real-time monitoring AFAIK.
- NathanKP 14y agoSame here. In fact, the AWS dashboard was still showing 2/2 checks passed for some 20 minutes after Pingdom told me my site was down. Then the AWS dashboard finally updated and told me that 3 minutes ago my instances became unreachable. That is pretty poor. AWS should be able to know right away and email me themselves.
- RegEx 14y agoI've learned to ignore the checks passed for quite a while, especially for servers on load balance.
- keithnoizu 14y agoBy over fifteen minutes in my case. Possibly thirty. WTH.
- stevefink 14y agoCpu0 : 0.3%us, 0.0%sy, 0.0%ni, 0.0%id, 99.7%wa, 0.0%hi, 0.0%si, 0.0%st <-- EBS subsystem is completely unreachable. I/O wait times are tanked across the board for me (I'm in US-EAST-1).
- nirvdrum 14y agoWhat zone? I really wish Amazon would provide that info, instead of saying that it only affects one zone.
- stevefink 14y agoBoth my MySQL master (I'm not using RDS) and Redis Master servers are affected and are located in zone us-east-1a.
- kokon 14y agoWhy do you have two masters in the same AZ?
- gabrtv 14y agoAFAIK zones are randomized. 1a for me is 1d for you.
- stevefink 14y agoThat's really interesting if it's true. I had never heard this before. Thanks for the tip.
- mattwdelong 14y agoDo you know why this is?
- lytfyre 14y agoHumans are predictable - it wouldn't spread the load across zones very well.
- rabble 14y agoGood time to consider Google's Compute Engine as an alternative? What will we call it, GCE?
- vachi 14y agoahh Acronyms
- jfoutz 14y agocurrently, it is a limited beta. Also, it looks to be more expensive.
- malachismith 14y agoActually, if you do the normalization to make it apples to apples (and adjust for the difference in RAM) it looks price competitive. My numbers make it look slightly more expensive than AWS EAST (teh suck) and slightly less expensive than AWS WEST.
- rabbitfang 14y agous-west-2 (Oregon) has identical pricing to us-east-1 (Virginia).
- zedwill 14y agoInteresting enough not only the EBS is down, but ELB can not register instances even if there are not EBS based and completely operational. I have some live instances running without EBS disks that I can not place behind the ELB as it is not working.
- oasisbob 14y agoI have some live instances running without EBS disks that I can not place behind the ELB as it is not working. ELBs are sometimes EBS backed.
- mikebo 14y agoWorst part of this outage: paying for a multi-az RDS instance and having failover totally, completely, fail.
- gouranga 14y agoThat sucks badly. Similar thing happened to me a while ago with a vendor. When your management team summons you to ask why the hell their site is down, you can't point fingers at the vendor if their marketing literature says it doesn't go down. Sticky situation.
- its_so_on 14y agoIf you don't host your data in several alternative dimensions so that the same events wouldn't transpire in all of them - why not assume you'll encounter the occasional outage?
- gouranga 14y agoIf only people understood that fact. Unfortunately few do.
- TazeTSchnitzel 14y agoCan't you tell management that it isn't as reliable as they claim?
- gouranga 14y agoI did. Unfortunately in the financial services industry, believing it means taking responsibility for it.
- keithnoizu 14y agoI'm paying like 2,300 a month and even something basic like failover isnt working. I'm not happy.
- 14y ago
- pearle 14y agoI'm running in us-east-1 and my EC2 instances and EBS volumes are still responding ok for the moment... Fingers crossed (just deployed to AWS less than 2 weeks ago).
- rdl 14y agoI'm curious why no public paas is multiple AWS region.
- malachismith 14y ago1) because AWS East is so much cheaper (and none of us like spending money) 2) AppFog actually is multi region (and multi IaaS as well)
- kanwisher 14y agoOregon is same as AWS East, seems to have a smaller set of boxes, have gotten errors in the past about not having any more servers to allocate.
- malachismith 14y agoSame in that they are both AWS and sometimes generate errors - yes. Not the same in that East has had four significant outages in the last 16 months and West has not.
- rdl 14y agoI'd tolerate multi-AZ as a baseline. Thanks for AppFog -- I hadn't heard of them, but will check them out.
- KenCochrane 14y ago9:32 AM PDT Connectivity has been restored to the affected subset of EC2 instances and EBS volumes in the single Availability Zone in the US-EAST-1 region. New instance launches are completing normally. Some of the affected EBS volumes are still re-mirroring causing increased IO latency for those volumes.
- bad_user 14y agoFor what is worth, my small website is online again.
- KenCochrane 14y agoI'm still seeing issues, some instances that aren't starting, and others I'm still not able to connect to. So I'm not sure what they are talking about.
- keithnoizu 14y agoI feel like you can't really say you're in the green when you still have customers unable to use your service. My instance is still stuck in failover. "9:39 AM PDT Networking connectivity has been restored to most of the affected RDS Database Instances in the single Availability Zone in the US-EAST-1 region. New instance launches are completing normally. We are continuing to work on restoring connectivity to the remaining affected RDS Database Instances."
- gooeyblob 14y agoAbsolutely agree - that's just silly. Their status page is close to useless.
- pwmanagerdied 14y agoAnybody else having trouble with S3? I don't have actual stats, but I'm noticing a lot of latency and poor performance starting recently, so I'm assuming it's related.
- tolos 14y agoEvery time (two out of two), by the time I click on "X is down" link, the service/website is working again. Surely there is a better platform for alerting about outages than ycombinator?
- bmelton 14y agoI was down for approximately three hours this morning. I don't know when this submission was posted, but I made one shortly after discovering the outage myself. Either way, if you're using RDS, even if this didn't affect you, it's discussion-worthy. I was affected, and we're building a not-yet-launched product that allows us the time to consider "Is Amazon really where we want to be?". The more failure I'm aware of, the more informed that decision is.
- pjscott 14y agoPingdom does a good job of it, if you point it at a public-facing web site you particularly care about. I'm not affiliated with them; I've just been woken up by them.
- pearle 14y agoAnyone have any details on why us-east-1 seems to be less reliable than the other regions? Is it the oldest?
- sausagefeet 14y agoI'm under the impression it's the most used.
- NoPiece 14y agoIt probably is the most used, being a cheaper alternative to us-west, but are you suggesting it fails more because it is used more? It does seem that the big AWS outages (in the us) have been concentrated in us-east. I have wondered if it just because us-east is newer so they haven't had has much time to work things out, or that the us-west team is a little better? edit: btw, I am not dismissing "used more" as a valid theory. More use = more hardware = more complexity which could lead to more failures.
- sausagefeet 14y agoMy theory is "used more".
- rabbitfang 14y agoThere are two different us-west regions. One in Oregon (priced the same as us-east) and one in California.
- malachismith 14y agoIt's the oldest, yes.
- jaylevitt 14y agoAccording to this calculation (which attempted to probe all the racks in EC2), over 70% of EC2 lives in us-east. http://huanliu.wordpress.com/2012/03/13/amazon-data-center-size/ http://huanliu.wordpress.com/2012/03/13/amazon-data-center-s...
- mattbillenstein 14y agoI suggest until Amazon uses RDS For their database - that you don't either...
- gregholmberg 14y agoIndividual availability zones can be identified using the API. ec2-describe-reserved-instances-offerings --region will tell you what the zone's identifier is. After you list the permanent identifiers, you can match them up to find out if your us-east-1a matches my -1d. This Alestic article shows how to label them all. [0] "Matching EC2 Availability Zones Across AWS Accounts" http://alestic.com/2009/07/ec2-availability-zones http://alestic.com/2009/07/ec2-availability-zones
- anuraj 14y agoMine is okay
- pjscott 14y agoEC2 comes with a free Chaos Monkey service. It's called EC2. I know, they're trying to make it reliable and they've got a bunch of very hard problems to solve. That doesn't change the fact that sometimes some of my servers just permanently stop responding to pings until you stop-start them, or get crazy-slow I/O, or get hit by these once-in-a-while-and-always-at-night outages. It's great when you suddenly need a hundred more servers, though.
- DigitalSea 14y agoIssue #3298392 for EC2 this month. This is ridiculous, so many websites rely on EC2 and it's proving to be extremely unreliable. Cloud computing is definitely not the answer to everything it would seem.