4 ms·
When this sort of thing happens, it's very easy to point out all the things you should've done differently. In retrospect, everything's easier. You can't antic
by jespern 17y ago
When this sort of thing happens, it's very easy to point out all the things you should've done differently. In retrospect, everything's easier.
You can't anticipate everything, and as I've pointed out in another comment here, this one is rather exotic.
Quick summary of what the problem is: We have an EBS volume. It mounts fine, appears fine. The problem is that it's excruciatingly slow. We can't serve data from the volume at any speed, really. Running an "ls" takes over a minute, in a small directory.
All systems are running, everything should be fine, but seeing as we can't read the data fast enough, we've been forced to put a static page explaining what's going on.
Booting a new instance, re-creating the volume from a recent snapshot, doesn't help. The exact same problem persists. Why? We don't know. Amazon's figuring it out.
We're doing everything we can do remedy the problem, but unfortunately right now, that consists of our team drinking coffee to not fall asleep, waiting for the final call from Amazon telling us they've sorted it out.
- gfodor 17y agoOk, that answers my other question. I think the fundamental issue (aside from the amazon issue) is you had bytes living on a single EBS disk that weren't replicated to another disk. For important data, this is probably a bad idea regardless of backup strategy, etc. Edit: By the way, the point here isn't to say "you guys screwed up" but to underscore that these types of issues aren't 100% Amazon's fault either, both parties had issues and the tone of your post seems to be "EC2 and EBS are not reliable we are switching off of it" when the truth lies somewhere in the middle.
- jespern 17y agoI apologize if that's the tone I'm relaying. I guess I'm just frustrated due to the time it's taking to fix it.
- gfodor 17y agoYea, it sounds like you guys are pretty much at Amazon's mercy right now, which sucks. I think a better way to look at these types of things is "what could we do to prevent this from happening again" and make a post about EBS gotchas. Since it sounds like you guys are talking about straight up file storage I'd guess a good option would be to set up an HDFS cluster and be smart about locality/replication to minimize the latency. Edit: Oh, and one more thing is take a backup that doesn't involve EBS snapshots. Maybe biweekly dump the entire sucker to S3 or something or have it getting pushed there all the time. This is something we've been meaning to do since the snapshotting capabilities of EBS are still a bit too magical for me to sleep well at night. (To be fair though, they've worked great when we've needed them to.)
- keefe 17y agoI don't think that this was a replication issue based on this comment : >Booting a new instance, re-creating the volume from a >recent snapshot, doesn't help. The exact same problem >persists. Why? We don't know. Amazon's figuring it out. If you can recreate the volume from a snapshot and hit it with fresh instances and run into the same problem, this is quite worrying. If it had resolved after a restore from backup, I would have felt better about EBS. As I see it there are only these options : 1) There is a general, systemic failure in EBS. You ran into it and highlighted it to AWS and they are fixing some problem. If other people are not having the same problem as you, I would be more inclined to think of #2. 2) Some usage pattern violates an assumption that was made when EBS was designed and screws it. Restoring from the backup reproduces the usage pattern. This could be simultaneous connections or # of distinct files in the volume, for example. One way to test this would be to split the data in the drive into a larger number of smaller EBS-es (EBSii? whatever the plural(: ) or throttle the simultaneous connections and see what happens. did I miss anything?
- idlewords 17y agoYes, the actual cause, a problem on the wire between them and Amazon. UDP flood eating up available bandwidth. I didn't guess it either :-)
- keefe 17y agoSo, do you think setting a smaller max size on your EBS instances would have avoided this by spreading the traffic, so if you were using 1TB using 10 100GB ones instead and federating queries across them?