12 ms·
Microsoft had three staff at Australian data centre campus when Azure went out
- collaborative 3y agoGuessing affected customers had to spend time and effort on top of ongoing high cloud bills I've slept so much better since I began hosting, producing energy, and cooling on-prem
- mattlondon 3y agoThe secret is hosting across failure boundaries so that a single outage like this does not impact you. Self-hosting is fine if you can afford the capex for two physically separate data centers (like really separate - like 100+ miles etc (or more!) to cope with natural disasters) and the staff to operate & maintain them 24/7. For many, this is not realistic. For those that do need to use cloud, just make sure you are running your services in different failure zones.
- gruez 3y ago>like 100+ miles etc (or more!) to cope with natural disasters) People talk about this often but this failure mode seems to never happen? When was the last time us-east-1 went down because of a natural calamity compared to some technical issue?
- mattlondon 3y agoNot sure about us-east-1 specifically but there are frequently fairly large natural disasters in the US - there are always hurricanes and stuff, there was that flooding in new York not so long ago, earthquakes in California in the 90s, wildfires etc. And this is just in the US. Basically, don't put all your servers in NYC or all in SF or whatever, but put half in NYC and half in SF and that random hurricane/wildfire/flood/snowstorm etc won't take out both of your data centers. .... Of course then you have latency issues to think about, but that is often quite application-specific and potentially a good problem to have if a slightly slow website or database or whatever is the biggest problem you have when the alternative would have been a total shutdown. There are also occasional fires and stuff that take out a whole building (I think OVH had this in France recently?). Ensure that your failure zones are physically separate places, and not just logically-separate zones in the same physical building, or in a building that is next to the one on fire :)
- gruez 3y ago>but there are frequently fairly large natural disasters in the US - there are always hurricanes and stuff, there was that flooding in new York not so long ago, earthquakes in California in the 90s, wildfires etc. Right but what type of datacenter related incidents did they cause? Did us-east-1 go down because of hurricane sandy? Did us-west-1 go down because of wildfires? I don't seem to remember any datacenter outages caused by wide area natural disasters, whereas I can remember plenty caused by BGP/DNS/config shenanigans.
- jquast 3y agoI remember Hurricane Katrina shutting down lots of online services, and directnic battling to stay online https://www.datacenterknowledge.com/archives/2007/11/05/provider-that-survived-katrina-exits-new-orleans https://www.datacenterknowledge.com/archives/2007/11/05/prov...
- TheNewsIsHere 3y ago> Did us-east-1 go Dow because of hurricane sandy? Nope, but Sandy did a hell of a lot of damage to some key telecommunications infrastructure. Verizon lost multiple floors worth of equipment, cabling, and related infrastructure that served at least their customers across Manhattan. Having geographical redundancy for mission critical workloads is a good investment if your business is making money. Networked computing is one of the few places we can actually “run away” from a physical source of problems. (Not forever, or universally, of course). We’re based on the eastern seaboard. You bet we have failsafes in areas less susceptible to natural disaster.
- toast0 3y ago> Did us-east-1 go down because of hurricane sandy? No, but I was at a company with all the production services in Reston, VA during that storm, and we would have been pretty screwed if Sandy made landfall in the DC area instead of continuing north. Sandy's flooding in NYC wasn't great for some of the datacenters there, I seem to recall some having trouble, but most were fine. BGP and DNS are certainly much better at causing disruption, and especially global disruption though.
- deleted 3y ago[deleted]
- traceroute66 3y ago> For those that do need to use cloud, just make sure you are running your services in different failure zones. By which time you might as well just roll out your own kit in colocation or your own datacentres. The cloud providers are nickle and dimers, they charge you for every little tiny thing. Cloud might look cheap at cents-per-hour, but then you find you need X "services" to deliver your Service and so you are talking about exponential cents-per-hour (X cloud services times x cents-per-hour). And then running your services across failure zones will of course cost you more beyond the basic double-cost, because most cloud providers charge by the GB for cross-zone traffic. So if you're doing cross-zone replication, that's gonna cost you a pretty penny. Meanwhile, in your own colo/DC, you have predictable costs. And you can get redundant connections between sites for a flat rate, not some stupid per GB fee.
- AugustoCAS 3y agoFully agree on this, plus (a very important plus) test that severing down an AZ doesn't bring the services on the good AZ down too. And test this frequently. I would be very, very surprised if the companies mentioned, in particular banks, weren't running on multiple AZs, but I wouldn't be surprised if the scenario of severing down an AZ was not tested.
- haolez 3y agoWhat about data center colocation? When you simply rent the energy, cooling, etc, but the hardware is yours? Do you think it's a nice middle ground?
- traceroute66 3y ago> Do you think it's a nice middle ground? It is. The cloud fanbois will tell you until their blue in the face that its not. I fully accept that the cloud is great for bursty workloads where you're doing nothing and then suddenly half the planet needs your service for a couple of days. That is clear. But if you've got a reasonably stable baseload running 24x7x365 and a few modest bursts here and there then honestly people need to do the math, because if you look at beyond the short-term figures, the cloud tends to work out much more expensive than colo if you look at for example a three-year period. Most people don't need the scale the cloud gives. They think they do, but really most people will never grow to FANG scale as much as they may dream it !
- dagw 3y agoI believe the real secret reason the cloud is so popular among developers (based on 10+ years of experience) is that cloud providers are so much nicer and faster to deal with than your company IT department. Also on the price side, I'm not comparing the price of cloud vs colo, but the price of cloud vs what the company IT department charges my department for being allowed to use one of 'their' colo servers, and that is many times what a cloud server costs. (as a real world example, the place I used to work internally invoiced $150/server/month for a virtual server that would cost me $20/server/month on AWS before any discounts). Cloud lives not by competing against smart people running their own servers, but against inefficient internal IT services, and there they have them beat both on price and quality.
- hotpotamus 3y agoI'm the ops guy on a small dev team, and I run a sort of hybrid setup for prod that does involve me working on hardware in a colo sometimes, though fairly rarely (I'd love to spend about half my time hauling servers around and cabling stuff so that I'm not stuck at a desk all day, but that's not the way it is). The whole point of my job is to enable developers to deliver code that provides customers value. On that level I actually embrace the common "condescensions" (so-to-speak) that I'm tech support for developers or a YAML wrangler. I actually had an experience recently where a developer asked to make some changes to our infrastructure. I pretty much developed our container orchestration system (based on Docker Swarm rather than Kubernetes - a choice our architect made that I've come to appreciate), so I walked him through how my IaC works, told him what he needed to change and then reviewed his pull request and applied the changes. I guess we're on a devops journey now if I want to put it in corpo-tech speak. Anyway, I suppose a lot of IT departments/guys get lost in creating their "perfect" unassailable systems and forget that the big picture is that the job is to enable customers; most directly are likely to be the developers or other internal employees, but ultimately the end customer who's handing you money to solve their problems.
- dgrin91 3y agoI know very little on the datacenter operations side of things - I guess 3 people is not a lot, but what is normal? How many operations people are at say AWS US-East-1? I presume it doesn't scale with number of servers, that would not scale well. What is a 'normal' level? 10? 100? It can't be more than 100, can it?
- dmw_ng 3y agoI doubt any amount of staffing could address lack of specialism in dealing with power or air conditioning issues, both likely involve infrequent maintenance by external vendors. 20 people blowing up the phone to a vendor doesn't fix a problem any faster than 1 etc.
- nixgeek 3y agoAmazon goes to the extreme of putting its own custom firmware on switchgear because the choices that vendor makes in theirs doesn’t align with their objectives. I don’t think AWS is blowing up a vendors phone when something goes wrong in one of their facilities. [1] https://www.datacenterknowledge.com/archives/2017/04/07/how-amazon-prevents-data-center-outages-like-deltas-150m-meltdown https://www.datacenterknowledge.com/archives/2017/04/07/how-...
- Scaevolus 3y agoAmazon doesn't make their own AC units or generators, so it's still likely they would need external support for a case like this.
- lithos 3y agoThat's some magical thinking, thinking that you don't need hardware people because you wrote some software. All the same gear is still there, its control functions are just ceded to Amazon panels. And integrated to the point of even removing some PLC like devices.
- mattlondon 3y agoI've heard of the big-cos having to use bonafide robots for doing manual tasks in a data center like replacing broken drives or swapping tapes etc. I think there is still a bunch of manual tasks to be done. That said I have no idea. When I worked (many many many many years ago) in a small DC that is perhaps the size of a 2bed apartment we had 4 guys scurrying about doing stuff (hands-on-keyboard, routing cables, replacing hardware etc). This was way before Docker & Kubernetes et al - physical iron and all that. I would assume that in modern DC ops you could run a football field sized DC with less than 10 people due to automation. But that said if part of the actual infrastructure like power or cooling fails, you need to have the right skill-set in place. If the cooler's failed and couldn't just be turned off and on again, we would have been out of luck in my old DC days and would need to call someone in and just hope the servers didn't fry in the meantime. Sounds like a similar deal here.
- WaitWaitWha 3y agoThree to five on-site staff to operate a mid-sized DC (10MW/~1K rack yield) is not unusual. This is assuming there are several others on-call.
- dijit 3y agoI guess people really have forgotten how to run datacenters. The only people who should be shocked in this thread are the people who have been hoodwinked into thinking operations is so hard you need thousands of staff. I know AWS/GCP/Azure like to charge us as if we were hiring an army of sysadmins, but the truth is that day-to-day DC ops does not require so many people. Hardware failures are more rare than you think and you can work around them without panicking anyway.
- benterix 3y agoThe management's way of thinking is: "Well, let's just pay for the peace of mind." Except that this famous peace of mind never comes, because the cloud gets more and more complex each year and it's hard to keep up. Heck, even Amazon can't keep up: for example, officially they depreciate bucket policies but internally they are using it for example in the Cloud Formation templates for the Control Tower. But now it's too late to go back as most of the internet is running on the three major public clouds. You need a lot of determination and a good plan to free oneself from vendor lock-in. In larger orgs it's practically impossible.
- adambatkin 3y agoI don't believe that S3 Bucket Policies are deprecated. They are powerful, effective, and consistent with almost everything else at AWS (Resource Policy). Perhaps you are thinking of ACLs?
- benterix 3y agoSorry, yes, I meant ACLs!
- wkat4242 3y agoThis peace of mind is also outsourcing responsibility. Having someone else to point to when shit hits the fan is very valuable for a manager. In this case they can't even get blamed for their vendor choice because both AWS and Azure are now so big that they're in "nobody ever got fired for buying IBM" territory.
- svaha1728 3y agoIt’s Microsoft. I’m kinda surprised it’s not ChatGPT-REPL at this point.
- PeterStuer 3y agoI wonder what a 4th, 5th or 10th person onsite could have done to speed up mitigation and recovery.
- AtNightWeCode 3y agoI straight out think MS lied when they said this was an .au only issue. We had a surge of rouge traffic from MS during this issue and we are pretty much on the other side of the globe.
- aaron695 3y ago[dead]