5 ms·
sorry, this is probably tangential, but how is the AWS outage any different from any other hardware outage, regardless of whether or not you actually own the ha
by unshift 15y ago
sorry, this is probably tangential, but how is the AWS outage any different from any other hardware outage, regardless of whether or not you actually own the hardware?
why does it seem that everyone expects 100% uptime from a VM just because it's in "the cloud"? shouldn't "the cloud" be used to make fault-tolerance even easier, because you have access to multiple geographic regions and multiple providers with little fuss?
they're still computers, and they're still bound to go down occasionally. i don't see how running leased VMs somehow absolves you of doing basic operations work and guarantees a bulletproof experience. regardless of where you host it, writing a nicely distributed and fault-tolerant system is and always has been difficult. the only part that's gotten easier is finding rack space.
- justinsb 15y agoWith physical hardware, the basic failure characteristics etc are well known. When the hardware is virtual or abstracted as it is on the cloud, you have no choice but to rely on the promises made by the provider. If you can't build a reliable system on the abstraction provided by the cloud, you can't reliably use the cloud. But with a good cloud, you can architect a reasonable solution based on the provider's promises. However, if the provider doesn't keep those promises, then all your hard work and calculations go totally out the window. At that point, you have to figure out a way to run a database on a system that effectively offers you no guarantees. That's what we're working on.
- dialtone 15y agoFor your application though a failure is a failure in any scenario. EBS went down in one zone and was slow in a couple others for a few hours, it's a bad downtime but some services survived the problem relatively unscathed. This suggests that something could have been done to avoid any trouble at all. If the datacenter you are in has a slight conditioning problem this summer and 10% of your drives breaks down due to excessive heat, how quickly will you be able to re-provision the data center? About a week after AWS outage the italian ISP Aruba had a UPS failure due to a fire in the UPS room and the entire datacenter switched off automatically during the night. For the following 8 hours that datacenter was off for every customer. How would your new solution handle such a situation? Designing a hot standby replication solution for MySQL/PostgreSQL that works across regions seems easier to me rather than implementing a database from scratch that should solve a very complex problem.
- justinsb 15y agoCertainly it is possible to engineer a hot standby solution for MySQL databases using DRBD, or synchronous replication etc. There's a choice between engineering an endless series of fixes like that, or re-examining the problem space and building a new database. After years of doing the former, we chose to examine the latter, and we're seeing good indications from that approach.
- jacques_chester 15y agoI have misgivings because you are not the first to try and re-examine the problem space. A lot of the endless series of fixes have arisen from just such attempts. I want you to succeed, but you're dealing with seriously hairy deep magic issues that the best minds in the industry and academia have been chipping away at for decades.
- justinsb 15y agoA healthy skepticism, then: nothing wrong with that. It's a speculative project with a big payoff if it succeeds, and I wouldn't ask you to believe until we've shown a working product. Your good wishes are appreciated though!
- unshift 15y agosorry but that seems really naive. an SLA is not a guarantee of service -- you can't force a computer or network to stay online because someone with a sheet of paper said so. it's a target for a best-effort guarantee and when that guarantee is broken then you get a bit of a refund. it's like when you have a big building project and there's a provision that the contractor will refund $5000/day for every day past march 1st. doesn't mean march 1st never gets passed. failure modes in the cloud are known too: your VMs are either working or they aren't, or maybe they're somehow degraded but you should fail over anyway. it's very similar to physical hardware. what will you think when a backhoe takes out your datacenter's fiber for 10 hours? that physical hardware and datacenters are now unreliable too? as for a database that runs on a system that offers no guarantees -- isn't that all of them?
- justinsb 15y agoTotally agree that SLAs are not a guarantee of service - I don't believe I said they were, and I was trying to make the exact same point: too many people treat them as if they were a guarantee, even when they carry only token penalty clauses. Separately from the SLA, technical promises/guarantees that AWS did make e.g. isolated AZs were broken in the April outage. I think that your proposed model ("machine is online or not") may be sufficiently simple that the AWS cloud can satisfy it; however I think it is very difficult to build anything interesting if that is the only axiom you have. In particular, I would want something related to persistent storage in the model, or else storing state becomes very difficult.