4 ms·
The duration of the downtime seems to have been dependent on not just disk throughput, but also network and CPU throughput. Eliot did mention that a large porti
by dirkgadsden 16y ago
The duration of the downtime seems to have been dependent on not just disk throughput, but also network and CPU throughput. Eliot did mention that a large portion of the downtime was partly caused by the slowness of EBS. Making a rough estimate, I think the downtime would probably be about 2/3 of how long it was if it had been on an SSD RAID. However, then you run into the issue of having to maintain your own servers; which, judging by the fact that they are using EC2 and EBS so heavily, Foursquare does not have any desire whatsoever to manage their own infrastructure.
- look_lookatme 16y agoIt's an interesting angle if that's the reason they are using EC2. Dedicated hosting might be more expensive, but hardware managed is pretty efficient at this point. Also, given the vast number of EC2 instances they could afford, it seems counter intuitive to be running only two mongrel shards. If you are that stingy about spreading data around, you might as well be using dedicated.
- dirkgadsden 16y agoIt doesn't seem that it was necessarily stinginess of neglect, but rather stinginess imposed by their situation. Like the article said, it took hours to create a new shard, downtime that Foursquare definitely did not want.
- look_lookatme 16y agoWhoops, I meant "mongo" shard!