5 ms·
Amazon EC2 down?
- deleted 14y ago[deleted]
- harryh 14y agoFWIW as of approx 0450 UTC we're starting to see various instances that had become unresponsive return back to service.
- ftwinnovations 14y agoYup, very down. My site is down, and I've been attempting to reboot and recreate instances like a madman... Why didn't I just check HN first??
- derekclapham 14y agoN. Virginia in my experience is by far the least reliable region on EC2/EBS... Fortunately our app servers are across 2 zone in the region... but our db server is just a lone master... Our slave is down... Very nervous.
- WALoeIII 14y agoI run a master with a "hot" master each in a different AZ and slaves of each in their respective AZs for days like today. Expensive, but makes it easy to sleep at night. The slaves have their EBS disks snapshotted every 30 minutes, the master every 24 hours.
- derekclapham 14y agoYeah... I snapshot our slaves every 30 minutes.. When you say "hot" master what exactly do you mean?
- WALoeIII 14y agoA "full-spec" machine, another X-Large with 4 EBS volumes that I can fail over to. Its in circular replication with the other "active" master (only one is receiving writes at a time). These instances are only snapshotted once a day to keep them as fast as possible.
- deleted 14y ago[deleted]
- dangrossman 14y agoN. Virginia is by far the most used region. There's more stuff to fail and any failure will affect more people. I don't think there's anything inherently less reliable about the geographical location of the building.
- csmcdermott 14y agoComing back up now...
- elq 14y agopower outage in one of their AZ's. Only after lots of machines died did the generator kick on. Lovely.
- jacobian 14y agoDo you have a source for this?
- elq 14y ago'tis what I heard on an outage call...
- robbiet480 14y ago"outage call" AWS provided, heroku provided or other? Need a good solution for this
- stevefink 14y agoAre you sure? I have a graph of one of my Resque boxes taking 30Gbit/s of inbound traffic. Looks like a DDoS attack to me.
- elq 14y agoThe update I heard was (essentially) 'Another update from Amazon: Looks like it was a power issue for one facility that services a particular AZ in us-east-1, flipped to generator, now back on power and in recovery mode.'
- chintan 14y agoWe learnt our lesson the hard way after the great AWScalypse of Apr 2011. The lesson: Use n>1 hosting companies (even if one of them promises a-z-multiregion-distributed-fault-tolerant-back-up)
- kunalmodi 14y agohow are you quickly failing over from one to another?
- monsenhor 14y agoWe work with 3 providers. Today 2 of then go down. Just Linode is ok now.
- kunalmodi 14y agoironically, we moved away from linode last month because they kept doing maintenance and turning off our machines without telling us
- citricsquid 14y agoreally? I've noticed an increase in maintenance with Linode recently but they've always been really great with notice, so much so I am consistently surprised. If there is ever going to be planned downtime I've got an email at least a week in advance and it gives me the option to migrate my linode to another server at my own convenience if their schedule isn't compatible with my needs. Do you not get these emails and options?
- kunalmodi 14y agoFor the scheduled ones, I think I did receive some prior notice, I might have gotten particularly unlucky with a bunch of emergency/network/etc. maintenance affecting many of my servers simultaneously.
- raverbashing 14y ago
- philip1209 14y agoHootsuite just reported that it is offline - I'm unsure if it is related. In other news, for once Reddit is working.
- monsenhor 14y agoOur instances still down. The AWS service health says: 9:27 PM PDT We continue to investigate this issue. We can confirm that there is both impact to volumes and instances in a single AZ in US-EAST-1 Region. We are also experiencing increased error rates and latencies on the EC2 APIs in the US-EAST-1 Region.
- cardmagic 14y agoThis is why http://AppFog.com/ http://AppFog.com/ is investing in multiple IaaS and is not being hit nearly as hard. You can still sign up and even create apps.
- Emouri 14y agoOn the other hand they don't have any prices listed, and their blog is down "Error establishing a database connection". This doesn't exactly inspire to go to them for hosting.
- Pythondj 14y agoand on the other hand, check out the http://status.appfog.com/ http://status.appfog.com/ page (hint: it doesn't exist)
- jswanson 14y agoNewest update: 10:29 PM PDT We can confirm a portion of a single Availability Zone in the US-EAST-1 Region lost power. We are actively restoring power to the effected EC2 instances and EBS volumes. We are continuing to see increased API errors. Customers might see increased errors trying to launch new instances in the Region. Source: http://status.aws.amazon.com/?rf http://status.aws.amazon.com/?rf Or: http://status.aws.amazon.com/rss/ec2-us-east-1.rss http://status.aws.amazon.com/rss/ec2-us-east-1.rss
- reustle 14y agoThey keep saying it happened in a single availability zone when I saw frantic tweets from people in 1A, 1C and 1D.
- adamlindsay 14y agoEveryone's A zone is different. So I could say A is down while someone else is saying B, and we could be talking about the exact same zone. It makes it difficult to say if it is more widespread or not.
- saurik 14y agoYou can match them up through a quirk in one of the APIs. http://alestic.com/2009/07/ec2-availability-zones http://alestic.com/2009/07/ec2-availability-zones My us-east-1a is the affected zone, which is 3a98bf7d-126d-411a-a612-3a57a62dc688 using the incantation on the site. (Oh, and to note: my us-east-1a was also the affected zone during the massive outage last year, and I believe I remember another outage sometime between then and now. I almost feel like every Amazon outage affects my zone. I kind of wonder if that availability zone just sucks ;P.)
- datr 14y agoMaybe I'm assuming wrong but I guess in your example zone A and B are the same and that the different zone names users see don't represent different ways of spreading resources. If so, why aren't they named consistently? If not, are there any details on how and why they've set up their zones (or am I overlooking another assumption I've made?)
- ww520 14y agoHmm, my sites are still up. They are at us-east-1d. Keeping fingers crossed.
- drivebyacct2 14y agoDoes anyone else find it strange that two Heroku posts made the frontpage considerably (in relative terms, obviously) earlier than "EC2 down"? I would think EC2 is a more common denominator for people, but maybe other hosts have better redundancy and thus there wasn't an immediate awareness? Or am I just overly curious and it's really just that some Heroku clients happened to notice before an at-large EC2 customer? edit: I don't mean to imply a conspiracy of some sort, upon a reread. I merely am curious if there are just that many Heroku users in particular on HN or somesuch?
- res0nat0r 14y agoIt is probably because of the large Heroku outage and post here just the other day, and people are trying to point out that they are down again as that is more dramatic than a normal AWS disruption.
- malachismith 14y agoa PaaS like Heroku ends up being "front line support" for AWS. if you use Heroku and your apps fail, you don't care if it is Amazon's fault - you blame Heroku
- yuvalo 14y agoStill having problems staring a few instances. We just started a campaign so i thought there were performance issues with our application so it took me a while to look for ec2 issues. sigh
- redditmigrant 14y agoI wonder if the power outage here has anything to do with this - http://www.dom.com/storm-center/dominion-electric-outage-map.jsp http://www.dom.com/storm-center/dominion-electric-outage-map...
- deleted 14y ago[deleted]
- dsirijus 14y agoAm I at fault in believing this happens drastically less in Europe datacenters, for any cloud service?
- yuvadam 14y agoI believe that, at least for AWS, the us-east data center is drastically larger than eu-west.
- malachismith 14y agoLarger, older, slower, more fragile. And cheaper.
- disbelief 14y agoI've been wondering this myself lately. It seems that every major EC2 outage hits US-East. By comparison, my US-West instances have way better uptime (granted, over a shorter test period). I've never tried the Europe or Asia zones, but I'm tempted to now.
- snorkel 14y agoIt does the same thing as OK.
- duwease 14y agoStill down here, over 12 hours at this point. This is probably the second time we've been hit with something on AWS in the last three months -- and you have to pay them to talk to someone about it. We're definitely moving to Linode ASAP..
- spartango 14y agoIf your application needs to be up constantly, then it should probably be at least multi-AZ scaled, if not multi-Region. Multi-AZ applications are not affected by this outage, and multi-AZ events are very rare. Living out of a single AZ is very risky.
- ShabbyDoo 14y agoSo, Amazon has said since the introduction of EC2 that, to ensure really high uptimes, customers should use multiple availability zones and architect their applications to survive an outage in a single availability zone. While I would question Amazon's competence if outages of any sort were overly frequent, Amazon has not had many at all and no recent cross-AZ ones. [This is correct, right?] I recognize that architecting applications to be performant across datacenters (tolerant of relatively high-latency replication), but Amazon seems to be a poster child for keeping its promises w.r.t. availability. Is my take on this incorrect?