7 ms·
I've built out many 42U racks in DC's in my time and there were a couple of rules that we never skipped: 1. Dual power in each server/device - One PSU was powe
by ItsBob 5y ago
I've built out many 42U racks in DC's in my time and there were a couple of rules that we never skipped:
1. Dual power in each server/device - One PSU was powered by one outlet, the other PSU by a different one with a different source meaning that we can lose a single power supply/circuit and nothing happens
2. Dual network (at minimum) - For the same reasons as above since the switches didn't always have dual power in them.
I've only had a DC fail once when the engineer was performing work on the power circuitry for the DC and thought he was taking down one, but was in fact the wrong one and took both power circuits down at the same time.
However, a power cut (in the traditional sense where the supplier has a failure so nothing comes in over the wire) should have literally zero effect!
What am I missing?
I've never worked anywhere with Amazon's budget so why are they not handling this? Is it more than just the imcoming supply being down?
- Bluecobra 5y ago> What am I missing? My guess is that they cheaped out in having redundant PSUs to get you to use multiple availability zones. (More zones = more revenue) Even a single PSU shouldn’t be an issue if they plugged in an ATS switch though.
- Godel_unicode 5y agoUnless the ATS breaks, which happens.
- mnordhoff 5y agoYup. I'm still upset (but not angry) about https://status.linode.com/incidents/kqhypy8v5cm8 https://status.linode.com/incidents/kqhypy8v5cm8.
- Bluecobra 5y agoFor sure, in my context I meant a ATS in single rack/cabinet. If that went bad the blast radius would be contained to a single cabinet. But yeah, anything can and will happen. At another place I worked at, a site UPS took down an entire server room. It was pretty nice Eaton system but there was some event that fried the whole thing. Eaton had to send an specialist to investigate the matter as those events are pretty rare.
- uluyol 5y agoWhy spend the cost on dual X and Y when you can failover to another cluster? For big DC workloads, it is usually, though not always, better to take the higher failure rate than add redundancy.
- ItsBob 5y agoReally? You'd think at Amazon's scale an additional PSU in a 1U custom-built server (I assume they're custom) would be a few tens of $ at most. Actually, now that I type that it makes sense. Scaling a few tens of dollars to a bajillion servers on the off-chance that you get an inbound power failure (quite rare I'd reckon) might cost more than what they'd lose if it does actually fail. So yeah, they're potentially just balancing the risk here and minimising cost on the hardware. Edit: changed grammar a bit.
- deleted 5y ago[deleted]
- vel0city 5y agoAt big cloud provider scale like Amazon, Azure, and Google they probably aren't even running PSUs at each server, they're probably doing DC at the rack these days. No point in having a million little transformers everywhere, far easier maintenance centralizing those and have multiple feeding the bus bars going to each rack.
- rainbowzootsuit 5y agoThe ones Im seeing designed have been moving the DC out to the cabinets with A/B 480VAC power feeds on the bus, and integrated DC inverters/rectifiers/batteries at the rack level. More modular and a lot less copper at 10x the voltage. Still a lot of copper.
- notyourday 5y ago> I've only had a DC fail once when the engineer was performing work on the power circuitry for the DC and thought he was taking down one, but was in fact the wrong one and took both power circuits down at the same time. This is all local scale. Your setup would not survive a data center scale power outage. At scale power outages are datacenter scale. Data centers lose supply lines. They lose transformers. Sometimes they lose primary feed and secondary feed at the same time. Automatic transfer switches cannot be tested periodically i.e. they are typically tested once. Testing them is not "fire up a generator and see if we can draw from it" It is cheaper to design a system that must be up which accounts for a data center being totally down and a portion of the system being totally unavailable than to add more datacenter mitigations.
- vel0city 5y agoThe only full datacenter outage I've personally experienced was a power maintenance tech testing the transfer switch between systems where the power was 90 degrees out of phase. Big oof.
- ItsBob 5y agoYes but if you have reliable power from two different sources then the biggest risk (I'd imagine) is the failover circuitry! Something that should be tested tbh. Also, there are banks of batteries and generators in between the power company cables and the kit: did they not kick-in? Again, this is all pure speculation: I have absolutely no idea of the exact failure, nor how their infrastructure is held together - this is all just speculation for the hell of it :)
- merlyn 5y agoFrying hardware can affect much wider scope. I've had bad power supplies fry out taking the whole power circuit with it, and thus half (or whatever fraction) of the rack's power. I've also had bad power supplies bring down the whole machine as they shunted everything internal too. When things go bad, anything can happen. You can provide the best effort, and it'll usually work as expected, but there will always be something that can overcome your best efforts.
- deleted 5y ago[deleted]
- growse 5y ago> 1. Dual power in each server/device - One PSU was powered by one outlet, the other PSU by a different one with a different source meaning that we can lose a single power supply/circuit and nothing happens Nothing happens if you remember that your new capacity limit per DC supply is 50% of the actual limit, and you're 100% confident that either of your supplies can seamlessly handle their load suddenly increasing by 100%. I've seen more than one failure in a DC where they wired it up as you described, had a whole power side fail, followed by the other side promptly also failing because it couldn't handle the sudden new load placed on it.
- dijit 5y agoEDIT: I misunderstood you were talking about power feeds, the normal case is the run "48% as if it's 100%" (because of power spikes, but also most types of transformers run more efficiently under specific levels of load (40-60). Normally this is factored into the Rack you buy from a hardware provider, they will tell you that you have 10A or 16A on each feed, if you exceed that: it will work, but you are overloading their feed and they might complain about it.
- praseodym 5y agoOP is talking about the DC power feed, not a single server PSU.
- dijit 5y agoYou don't get fed DC power, you get fed AC power. But, point taken: yes your power feed should be running at <50%. But that just means you treat 50% as 100% just like any resource. Mostly this is outsourced to the datacenter provider; they'll give you a per side rating. (usually 10A or 16A) which also matches the cooling profile of the cabinet.
- vel0city 5y agoI mean, in some datacenters they run DC power to each rack. Its definitely more esoteric than having each device run AC but some people do it. However, with their comment DC == Data Center, not Direct Current.
- lordnacho 5y agoWhat about a UPS/battery thingy? That's saved me a few times, though it normally just gives enough time for a short outage. Is it uncommon in cloud infra?
- vel0city 5y agoFor even regular datacenters they'll often have UPS systems the size of a small car, usually several of these, to power the entire datacenters for a few minutes to get the diesel generator started.
- bob1029 5y ago> I've never worked anywhere with Amazon's budget so why are they not handling this? Perhaps we are going to discover how AWS produces such lofty margins by way of their next RCA publication.
- deleted 5y ago[deleted]