5 ms·
I feel like an idiot. MS featured my Azure startup today, quoting me about overall stability etc (which has been the case for us, until today). They then proc
by photorized 12y ago
I feel like an idiot. MS featured my Azure startup today, quoting me about overall stability etc (which has been the case for us, until today). They then proceeded to go down, taking all our production systems with them.
(yes we do have AWS, too)
Sigh.
- photorized 12y agoMicrosoft, you've got to be kidding me. Just tried opening a billing ticket, completed the forms in detail, attached screenshots, clicked Submit... 'Unable to Submit Request We are unable to complete the incident submission process at this time. Please refer to this page for phone numbers to call for Azure support.'
- Encosia 12y agoFor what it's worth, I've been watching several Azure-hosted sites that I control and they've been coming back online sequentially (and are all back online now). Whatever they're fixing, it seems to be taking some time, but is progressing steadily at a good pace in the last hour.
- photorized 12y agoMine have been gradually coming back, too. Timing couldn't have been any better for me. Some Alanis material there: https://twitter.com/bizspark/status/534858596748906496 https://twitter.com/bizspark/status/534858596748906496 Now I'll have to distribute between AWS and Azure, too.
- toomuchtodo 12y agoWhy not stick with AWS? (Just curious!)
- photorized 12y agoThey all have outages periodically, in my experience.
- Encosia 12y agoYeah, that's the main thing to take away from these events. I've never seen a system (cloud, local, co-located, or otherwise) with 100% uptime, despite every effort to the contrary. Even sites like Facebook and Google have downtime. In the last few years that I've been using it, Azure has been at least as stable as other cloud providers. That doesn't help with the awful timing though. Ouch. I just Buffer-retweeted your BizSpark tweet above, scheduled for tomorrow. Maybe a little bump now that you're back up and running will help ease the pain...
- icantthinkofone 12y agoFor nine hours?
- photorized 12y agoThank you.
- osipov 12y agoCheck out Softlayer. It has been more reliable than AWS in my experience.
- toomuchtodo 12y agoReally? I manage several hundred VMs and their associated EBS volumes on AWS and we've had 0 problems. Also no problems with S3. Ever.
- razzberryman 12y agoThey should run their support system on AWS since they're likely to get a lot of ticket requests if Azure is down. :)
- cddotdotslash 12y agoJust curious - if you have AWS too, then why did it take everything down? Can't you just swap the DNS?
- philwelch 12y agoDoes DNS propagate quickly enough to alleviate an outage or is it just a matter of ensuring that you recover within a few hours rather on waiting on an outage resolution that might take longer? Alternately, can't you just have multiple A records to distribute your load across cloud platforms and just drop the one for whichever platform is having an outage?
- latch 12y agoFrom experience with multi-datacenter setups, if you set a 60second TTL on your DNS records, you'll see 95%+ of traffic get the update within 5 minutes. Also, you can associate multiple addresses with a record. It's up to the client to retry on failure, but all browsers do (as far as I know)
- cube00 12y agoWouldn't that kill DNS if everyone did that considering it relies on caching for performance across the world?
- latch 12y agoI'm no expert, but no. Most big sites rely on a fairly short TTL. It's a thick layer of caches. Your browser, OS, router, ISP, and a bunch of intermediaries can cache the DNS. So even at 60s, you get good cache hits (the busier, the more true that is, of course) Also, the update can always happen asynchronously. You and 9999 people ask your ISP for Facebook's IP. It serves all of you a slightly stale IP and asynchronously fetches a new one (thus turning 10000 requests into 1). AKA: thundering heard problem. DNS mostly uses UDP, which is more efficient for the server and harder to DOS (the server doesn't have to maintain state per request). Finally, # of requests is usually (always?) a factor in the price of DNS services. So the cost is borne by the clients, not the service providers. And since DNS hosting is seemingly profitable, I assume they're more than happy to build up the infrastructure to deal with additional requests.
- higherpurpose 12y agoAzure has had quite a few outages this year. I'd say it's already lower than that 99.999 percent uptime or w/e they are advertising.
- rbanffy 12y agoAccording to https://cloudharmony.com/status-1year-for-azure https://cloudharmony.com/status-1year-for-azure, they didn't reach 99.99%.
- toyg 12y ago0.001% of 365 days is 8.76 hours. So yeah, shot for the year; but of course they'll do some "hollywood timekeeping" (or just ignore the matter altogether) and keep advertising...
- kazoolist 12y ago0.001% of 365 days = 5.26 minutes (http://www.wolframalpha.com/input/?i=0.001%25+of+365+days http://www.wolframalpha.com/input/?i=0.001%25+of+365+days)
- cddotdotslash 12y agoAnd the year isn't even over yet!