3 ms·
Does anyone else think that Reddit's usage of EBS might be the culprit? Looking at their past outage response: http://blog.reddit.com/2010/01/why-did-we-take-
by c2 15y ago
Does anyone else think that Reddit's usage of EBS might be the culprit?
Looking at their past outage response:
http://blog.reddit.com/2010/01/why-did-we-take-reddit-down-for-71.html http://blog.reddit.com/2010/01/why-did-we-take-reddit-down-f...
Money quote: "In response, we started upgrading some of our databases to use a software RAID of EBS disks, which gives drastically increased performance (at a higher cost of course)."
RAIDing EBS disks seems like a really really BAD idea. There is a non-trivial failure rate of any single EBS disk, and if you RAID them together, your failure rate of the RAID will consequently increase. Am I understanding that correctly?
If they fix that, could that be a 'silver bullet' to fix these outages?
- magicofpi 15y agoThey certainly think there's a problem with EBS: "Since that last failure, we have been doing everything we can to move ourselves off of the EBS product. We're about half way there. All of our Cassandra nodes are now using only local disk, and we hope to have all of postgres on local disk soon."