7 ms·
How come they dont have power backups?
by GrumpyNl 5y ago
How come they dont have power backups?
- trelane 5y agoAnything can fail, even your backup, and especially if it's mechanical.
- rdines 5y agoThe battery backups (called uninterruptible power supplies) are only meant to bridge the gap between the power going out and the generator turning on, which is a few minutes. Did they say power was the issue this time? I suspect it’s actually something else (ahem network)
- chkhd 5y ago"When a fail-safe system fails, it fails by failing to fail-safe." - https://en.wikipedia.org/wiki/Systemantics https://en.wikipedia.org/wiki/Systemantics
- 2-718-281-828 5y agois that just playing with words?
- itsoktocry 5y ago>is that just playing with words? It conveys reality, that "fail-safe" isn't literal, as if anyone believed that.
- 2-718-281-828 5y agoI mean it has to be play with words or tongue in cheek simply b/c the assumption of a fail-safe system failing is already contradictory. So you cannot say anything smart about that beyond - there are no fail-safe systems that fail.
- seeking_future 5y agoThe real world is the play. Words are just catching up.
- frupert52 5y agoDo you mean in that it fails by failing to be the thing that it purports to be? Making it no longer that thing? At what point does bread become toast?
- Talanes 5y agohttps://en.wikipedia.org/wiki/Gare_de_Lyon_rail_accident https://en.wikipedia.org/wiki/Gare_de_Lyon_rail_accident Fail safes do fail. Often due to severe user error.
- the-dude 5y agoAn unknown unknown.
- marcosdumay 5y agoYou mean to ask if it's a joke? Yes, it's a joke. Or you ask if it's a lesson about how real systems operate? Because yes, it's a very serious lesson about how real systems operate. Anyway, you seem out of grasp on system engineering. Your reply downthread isn't applicable (of course fail-safes can fail, anything can fail). If you want to learn more on this area (not everybody wants, and its ok), following that link of system theory books on the wiki may be a good idea. Or maybe start at the root: https://en.wikipedia.org/wiki/Systems_theory https://en.wikipedia.org/wiki/Systems_theory Notice that there is a huge amount of handwaving in system engineering. I don't think this is good, but I don't think it's avoidable either.
- jerf 5y ago"Notice that there is a huge amount of handwaving in system engineering. I don't think this is good, but I don't think it's avoidable either." In my experience, you can be specific, but then you get the problem that people think that if they just 'what if' a narrow solution to the particular problem you're presenting they've invalidated the example, when the point was 1. that this is a representative problem, not this specific problem and 2. in real life you don't get a big arrow pointing at the exact problem 3. in real life you don't have one of these problems, your entire system is made out of these problems, because you can't help but have them, and 4. availability bias: the fact that I'm pointing an arrow at this problem for demonstration purposes makes it very easy to see, but in real life, you wouldn't have a guarantee that the problem you see is the most important one. There's a certain mindset that can only be acquired through experience. Then you can talk systems engineering to other systems engineers and it makes sense. But prior to that it just sounds like people making excuses or telling silly stories or something. "(of course fail-safes can fail, anything can fail)" Another way to think of it is the correlation between failure. In principle, you want all your failures to be uncorrelated, so you can do analysis assuming they're all independent events, which means you can use high school statistics on them. Unfortunately, in real life there's a long tail (but a completely real tail) of correlation you can't get rid of. If nothing else, things are physically correlated by virtue of existing in the same physical location... if a server catches fire, you're going to experience all sorts of highly correlated failures in that location. And "just don't let things catch fire" isn't terribly practical, unfortunately. Which reiterates the theme that in real life, you generally have very incomplete data to be operating on. I don't have a machine that I can take into my data center and point at my servers and get a "fire will start in this server in 89 hours" readout. I don't get a heads up that the world's largest DDOS is about to be fired at my system in ten minutes. I don't get a heads up that a catastrophic security vulnerability is about to come out in the largest logging library for the largest language and I'm going to have a never-before-seen random rolling restart on half the services in my company with who knows what consequences. All the little sample problems I can give in order to demonstrate systems engineering problems imply a degree of visibility you don't get in real life.
- losvedir 5y agoNo. For example train signalling which controls whether a train can do onto a section of track operates in a fail safe manner, where if something goes wrong, the signal fails into a safe "closed" state rather than an unsafe "open" state. This means trains are incorrectly being told to stop even though technically the tracks are clear, rather than incorrectly being told to go even though there is another train ahead. "fail-safe" doesn't mean "doesn't fail", it means that the failure mode chooses false negatives or false positives (depending on the context) to be on the safe side.
- NovemberWhiskey 5y agoI think it's predicated on a misunderstanding of what "fail-safe" actually means. For example, in railway signaling, drivers are trained to interpret a signal with no light as the most restrictive aspect (e.g. "danger"). That way, any failure of a bulb in a colored light signal, or a failure of the signal as a whole, results in a safe outcome (albeit that the train might be delayed while the driver calls up the signaler). Or, in another example from the railways, the air brake system on a train is configured such that a loss of air pressure causes emergency brake activation. Fail-safe doesn't mean "able to continue operation in the presence of failures"; it means "systematically safe in the presence of failure". Systems which require "liveness" (e.g. fly-by-wire for a relaxed stability aircraft) need different safety mechanisms because failure of the control law is never safe.
- jsmith99 5y agoOr nuclear reactors that fail safe by dropping all the control rods into the core to stop all activity. The reactor may be permanently ruined after that (with a cost of hundreds of millions or billions to revert) but there will be no risk of meltdown.
- NovemberWhiskey 5y agoI don't know enough about reactor control systems to be sure on that one. The idea of a fail-safe system is not that there's an easy way to shut them down, but more that the ways we expect the component parts of a system to fail result in the safe state. e.g. consider a railway track circuit - this is the way that a signaling system knows whether a particular block of a track is occupied by a train or not. The wheels and axle are conductive so you can measure this electrically by determining whether there's a circuit between the rails or not. The naive way to do this would be to say something like "OK, we'll apply a voltage to one rail, and if we see a current flowing between the rails we'll say the block is occupied." This is not fail-safe. Say the rail has a small break, or if power is interrupted: no current will flow, so the track always looks unoccupied even if there's a train. The better way is to say "We'll apply a voltage to one rail, but we'll have the rails connected together in a circuit during normal operation. That will energize a relay which will cause the track to indicate clear. If a train is on the track, then we'll get a short circuit, which will cause the relay to de-energize, indicating the track is occupied." If the power fails, it shows the track occupied because the relay opens. If the rail develops a crack, the circuit opens, again causing the relay to open and indicate the track is occupied. If the relay fails, then as long as it fails open (which is the predominant failure mode of relays) the track is also indicated as occupied.
- thetinguy 5y agoThey do. I remember watching one of their sessions where they showed every rack having its own battery backup.
- tyingq 5y agoAn article on that: https://datacenterfrontier.com/aws-designs-in-rack-micro-ups-units-for-a-more-efficient-cloud/ https://datacenterfrontier.com/aws-designs-in-rack-micro-ups... Interesting quote: “This is exactly the sort of design that lets me sleep like a baby,” said DeSantis. “And indeed, this new design is getting even better availability” – better than “seven nines” or 99.99999 percent uptime, DeSantis said.
- taf2 5y agoit was not a total power loss. out of 40 instances we had running at the time of the incident only 5 of our instances appeared to be lost to the power outage. the bigger issue for us was ec2 api to stop/start these instances appeared to be unavailable (but probably due to the rack these instances were in having no power). The other issue that was impactful to us was that many of the remaining running instances in the zone had intermittent connectivity out to the internet. Additionally, the incident was made worse by many of our supporting vendors being impacted as well... IMO it was handled rather well and fast by AWS... not saying we shouldn't beat them up (for a discount) but being honest this wasn't that bad.
- res0nat0r 5y agoIf the rack your instances are running in are totally offline then the ec2 api unfortunately can't talk to the dom0 and tell the instances to stop/start, so you get annoying "stuck instances", and really can't do anything until the rack is back online and able to respond to API calls unfortunately.
- TrueDuality 5y agoAccording to the SOC certifications they give their customers they do.
- redm 5y agoSome datacenter failures aren't related to redundancy. Some examples: 1) transfer switch failure where you can't switch over to backup generators and the UPS runs out, 2) someone accidentally hits the EOD, 3) maintenance work makes a mistake such as turning off the wrong circuits, 4) cooling doesn't switch over fully to backups and while your systems have power, its too hot to run. The list can go on and on. I'm not sure why this is a big deal though, this is why Amazon has multiple AZ's. If your in one AZ, you take your chances.
- Spooky23 5y agoTheir datacenter(s) aren’t magic because they are AWS. That facility is probably a decade old and like anything else as it ages the technical and maintenance debt makes management more challenging.
- chousuke 5y agoSometimes, you have a component which fails in such a way that your redundancies can't really help. I once had to prepare for a total blackout scenario in a datacenter because there was a fault in the power supply system that required bypassing major systems to fix. Had some mistake or fault happened during those critical moments, all power would've been lost. Well-designed redundancy makes high-impact incidents less likely, but you're not immune to Murphy's law.
- macintux 5y agoTo my mind, among the more frustrating aspects to implementing protection against failure is that the mechanisms to be added can themselves cause failure. It's turtles all the way down.
- chousuke 5y agoYou need to pick your battles and choose what you want to protect against to mitigate risk and enable day-to-day operations. For example, too often people will set up clustered databases and whatnot because "they need HA" without much thought about all the other potential effects of using a cluster, such as much more complicated recovery scenarios. In the vast majority of cases, an active-passive replicated database with manual failover is likely to have fewer pitfalls and gives you the same operational HA a clustered database would, even though in the case of a (rare) real failure it would not automatically recover like a cluster might.