5 ms·
Amazon EC2 Outage Takes Down Foursquare, Instagram, Quora, Reddit, Etc
- scrod 15y agoLOL, the cloud.
- deleted 15y ago[deleted]
- Vexenon 15y agoMy reaction: shocking (but not really).
- bane 15y agoWell there goes all the parts of the Internet I'm interested in. Time to go read a book.
- sibsibsib 15y agoreddit is currently working for me.
- i386 15y agoand here I was thinking the change I just rolled out to our EC2 instances had boned our test environment. Two failures in a week? Does not really inspire confidence right now :(
- stevenp 15y agoMy t1.micro instance in us-east-1b seems to be up and running just fine as far as I can tell.
- thechut 15y agoThis may be a stretch...but anything to do with the Verizon line workers strike?
- fizx 15y agoIt's back now.
- fourspace 15y agoMaybe I'm oversimplifying things, but why haven't these companies distributed their compute resources across various facilities and cloud providers, enabled instant failover, and tested this before outages like these?
- talonx 15y agoThey will, now :)
- protagonist_h 15y agoThis is not an easy thing to do, especially if you use EC2 in conjunction with EBS volumes, which is typical setup. EBS volumes are created in a particular availability zone ("facility") and can only be accessed from the same zone. Therefore you not only need to distribute computing resources but also data, which is significantly harder. So even distributing across multiple Amazon data centers is not that simple. To distribute across different cloud providers you would have to rewrite large chucks of code for each provider or come up with some way to "abstract away" cloud providers. Either way it would be nightmare to manage and is likely not worth it for a typical startup. But yes, this CAN be done.
- smanek 15y agoIt costs engineering time to do so. Time that could otherwise be used to build features, better protect against more common failures, attract users, etc. Amazon probably has ~5hrs/year of complete failure of a region. Figure, conservatively, it would take 3 months of engineering time to protect against that, plus a 'continuing' cost of 1/2 a week per month to maintain that protection. You'd also have to (at least) double your provisioned capacity (which may include a larger ops team, etc). Assuming your servers cost $20k/month and devs cost $100/hr (both fully loaded), we're talking about ~$340,000 to prevent 5 hours of downtime (just for the first year). If downtime costs you more than $50K/hr, then it might make sense to be that fault tolerant. Otherwise, there might be better places for a startup to spend its (limited) resources.
- cageface 15y agoNot to mention that it's very easy to increase overall downtime by introducing all the extra complexity this kind of redundancy can bring.
- protagonist_h 15y agowe were thinking to migrate our service to EC2 from our dedicated softlayer server. now we will probably stick with our current setup.
- lars512 15y agoIt's a testament to how successful Amazon has been in its cloud offering. We're used to sites going down for one reason or another. What's weird is that Amazon's success has made all these failures so correlated. It's a strange feeling when many sites you like all fail at once.
- dbuizert 15y agoSomeone said business continuity? It can be costly, but could save your business. Stop saving that VC money and start saving your business.