3 ms·
I don't necessarily agree with your first two sentences, but I definitely agree with the last one. I know they're smart. I do have some concerns that they're h
by blaines 14y ago
I don't necessarily agree with your first two sentences, but I definitely agree with the last one. I know they're smart.
I do have some concerns that they're having too much downtime. If there's one small flaw in the system it seems that the whole thing begins to fail.
If they fix the problem, and it impacts the larger system in some other unknown way, a different equally crippling issue could present itself in the future. I'd like to be sure they're putting a huge effort into making sure these problems don't happen, and I don't have those assurances at the moment.
First things first, a better status dashboard that actually reflects how issues impact customers is needed. I'd rather have everything working fine and the status be 'red' than have servers down, support tickets, calls, emails, etc and see a 'green' on the dashboard.
- sokoloff 14y agoIf there's one small flaw in the system it seems that the whole thing begins to fail. I wonder if there is selection bias underlying that judgment? Reading all of Amazon's post-mortems of big events, it does seem that the whole thing is fragile. What I suspect is more likely true is that AWS suffers thousands of small failures every month and most are contained as designed, with no(or minuscule) customer impact. It's the ones that turn into highly visible failures that we read about. That said, I agree with you that EBS in particular seems to have more downtime than I'd expect. (And that other services like ELB depends on it makes it cascade in a way that's hard to design highly-available systems.)
- Firehed 14y ago> What I suspect is more likely true is that AWS suffers thousands of small failures every month and most are contained as designed, with no(or minuscule) customer impact. Isn't that the whole point of moving to The Cloud? There's supposed to be some magical system in place such that hardware failures are routed around and don't interrupt service. Of course you can roll this yourself with your own hardware, but this is done for you. It should be no small surprise that a system complicated enough to appear magical has some crazy complexity behind the scenes, and accidental dependencies can result in catastrophic failure.
- mbesto 14y agoDo you think maybe this is just a focusing illusion?[1] And therefore the utility you associate with the service is not correctly attributed. [1] http://en.wikipedia.org/wiki/Anchoring http://en.wikipedia.org/wiki/Anchoring
- forensic 14y agoI like this idea. But which aspect of his belief do you think might be an illusion? I'm not clear precisely what you're referring to.
- gizmo686 14y ago>I do have some concerns that they're having too much downtime. If there's one small flaw in the system it seems that the whole thing begins to fail. I think this is a case of selection bias. Most of the time, when their is a small flaw, the system transparently bypasses the flaw and the service continues uninterrupted while they fix the initial problem. Because of this, the only failures that people see are the ones where the bypassing process fails, in which case the problem affects many people. From the point of view of a single service running off of AWS, this is a much more stable system, becuase it will provide your service much more than a system without the auto bypass infastructure. From a end-user point of view, this benefit is not so clear, because while any given service is more reliable, they tend to go down at the same time.