3 ms·
So far, for last 8 months, we have seen couple of SMPS failures and a motherboard. Also had issues with UPS batteries. We are learning and improving in componen
by khadim 14y ago
So far, for last 8 months, we have seen couple of SMPS failures and a motherboard. Also had issues with UPS batteries. We are learning and improving in component selections, which should further improve
Failover is also managed at application layer by FOSS such as Cassandra, Hadoop, so user's have in general not faced much downtime.
If we are not able to scale or manage, we will re-look and consider moving to public cloud.
- peterwwillis 14y agoI'd wager the biggest cause of failures is improper air flow. Even if you keep the room ice cold, if the servers can't push air efficiently over the HDDs/CPU and into a hot aisle, you'll get hot spots and hardware failure is inevitable.
- khadim 14y agoThanks Peter. We are getting good pointers. Will try to figure out some way to fix.