5 ms·
I'm here to answer questions if there are any (I run Bitbucket.)
by jespern 17y ago
I'm here to answer questions if there are any (I run Bitbucket.)
- shizcakes 17y agoWhat are you thinking in terms of movement / failover at this point?
- jespern 17y agoWe've been contacted by several hosting providers, and right now, Rackspace seems pretty nice. The problem isn't that we don't have failover here, it's that we store all repositories on a single EBS volume. This has worked great for us in the past, but as of last night, that volume has become virtually unavailable to us. It doesn't matter which instance we mount it on, the throughput we get from it is excruciating. If, or at this point--when, we move, the disk architecture will look different, and general failover will be less of an issue. Amazon has for the past 8-10 hours been investigating the issue, and we're left pretty dumbfounded as of to what has happened exactly. I'll summarize everything in a blog post once the chaos is over.
- keefe 17y agoWhat do you think of the other post critcizing your architecture? I am doing my own arch work on EC2, so I am trying to understand exactly what caused this failure. Is this a problem with your instance being able to access any EBS? Why couldn't you spin up another instance with a fresh EBS from a backup and redirect DNS to that instance?
- jespern 17y agoWhen this sort of thing happens, it's very easy to point out all the things you should've done differently. In retrospect, everything's easier. You can't anticipate everything, and as I've pointed out in another comment here, this one is rather exotic. Quick summary of what the problem is: We have an EBS volume. It mounts fine, appears fine. The problem is that it's excruciatingly slow. We can't serve data from the volume at any speed, really. Running an "ls" takes over a minute, in a small directory. All systems are running, everything should be fine, but seeing as we can't read the data fast enough, we've been forced to put a static page explaining what's going on. Booting a new instance, re-creating the volume from a recent snapshot, doesn't help. The exact same problem persists. Why? We don't know. Amazon's figuring it out. We're doing everything we can do remedy the problem, but unfortunately right now, that consists of our team drinking coffee to not fall asleep, waiting for the final call from Amazon telling us they've sorted it out.
- gfodor 17y agoOk, that answers my other question. I think the fundamental issue (aside from the amazon issue) is you had bytes living on a single EBS disk that weren't replicated to another disk. For important data, this is probably a bad idea regardless of backup strategy, etc. Edit: By the way, the point here isn't to say "you guys screwed up" but to underscore that these types of issues aren't 100% Amazon's fault either, both parties had issues and the tone of your post seems to be "EC2 and EBS are not reliable we are switching off of it" when the truth lies somewhere in the middle.
- jespern 17y agoI apologize if that's the tone I'm relaying. I guess I'm just frustrated due to the time it's taking to fix it.
- gfodor 17y agoYea, it sounds like you guys are pretty much at Amazon's mercy right now, which sucks. I think a better way to look at these types of things is "what could we do to prevent this from happening again" and make a post about EBS gotchas. Since it sounds like you guys are talking about straight up file storage I'd guess a good option would be to set up an HDFS cluster and be smart about locality/replication to minimize the latency. Edit: Oh, and one more thing is take a backup that doesn't involve EBS snapshots. Maybe biweekly dump the entire sucker to S3 or something or have it getting pushed there all the time. This is something we've been meaning to do since the snapshotting capabilities of EBS are still a bit too magical for me to sleep well at night. (To be fair though, they've worked great when we've needed them to.)
- keefe 17y agoI don't think that this was a replication issue based on this comment : >Booting a new instance, re-creating the volume from a >recent snapshot, doesn't help. The exact same problem >persists. Why? We don't know. Amazon's figuring it out. If you can recreate the volume from a snapshot and hit it with fresh instances and run into the same problem, this is quite worrying. If it had resolved after a restore from backup, I would have felt better about EBS. As I see it there are only these options : 1) There is a general, systemic failure in EBS. You ran into it and highlighted it to AWS and they are fixing some problem. If other people are not having the same problem as you, I would be more inclined to think of #2. 2) Some usage pattern violates an assumption that was made when EBS was designed and screws it. Restoring from the backup reproduces the usage pattern. This could be simultaneous connections or # of distinct files in the volume, for example. One way to test this would be to split the data in the drive into a larger number of smaller EBS-es (EBSii? whatever the plural(: ) or throttle the simultaneous connections and see what happens. did I miss anything?
- pvg 17y agoSite seems to be back up, any better idea what happened?
- jespern 17y agoIt was fixed around 4am (GMT+2) last night, with the assistance of Amazon. I'm just going to summarize what happened here: We were attacked. Massive UDP DDOS. The flood of traffic prevented us from accessing our EBS store with any acceptable speeds, which is what caused everyone to think the problem was between our EC2 and the EBS. Of course this also explains why booting up a new instance and EBS didn't help anything. Also, it's happening again now, and we're working with Amazon to remedy it once more.
- tlrobinson 17y agoIs there anything Amazon could have done to prevent this (or at least made diagnosing it easier), or is it a problem with your particular application?
- jespern 17y agoWe're talking UDP flood here, saturating our bandwidth. It never reached our servers, it just ate all the bandwidth on our connection. I guess what Amazon could have done is be quicker in spotting the DDOS and take measures to prevent it.
- spudlyo 17y agoSo you never saw any evidence of this DDOS yourself? I'm somewhat skeptical of this explanation. It seems to me with shared infrastructure it'd be difficult to saturate just one customer's connection. It also doesn't make sense to me that this could be done without the traffic ever reaching your server. You used the phrases "our bandwidth" and "our connection" do things really work this way on the AWS cloud? Anyway, I'm really sorry you guys had to go through all of this, and I hope whatever it is that caused it is fixed.
- tlrobinson 17y agoSo it was actually entirely unrelated to EBS? The reason it was taking 10 seconds to do an "ls" was simply a saturated connection to your server, not too much EBS activity?