5 ms·
So they didn't monitor temperature of all systems by default to catch cooling problems? Sounds more embarrassing to me. That's fairly basic mistakes.
by nn3 7y ago
So they didn't monitor temperature of all systems by default to catch cooling problems?
Sounds more embarrassing to me.
That's fairly basic mistakes.
- gnarbarian 7y agoHindsight is 20/20. And there is at least some sort of survivership bias. In a huge system with lots of failover you won't notice all the contingencies they identified planed for and successfully mitigated. Then something breaks which; you thought would be covered under another redundancy, was covered until a recent change (firmware change to the HVAC system), you didn't plan for etc... And everyone points at the failure and says "that seems obvious", meanwhile the mountain of tests and monitoring and redundancy goes unnoticed.
- alexandercrohde 7y agoWell, to point, at the very least the SRE post-mortem should have said "And we considered distributing monitoring software to all our datacenters for future alerting" [Not to say that every company out there would do this, nor that this wasn't necessarily considered, but definitely questioning the merit/objectivity of a brag-piece that is trying to rebuild the waning google hype]
- Dylan16807 7y agoNah, "monitor temperature" isn't a hindsight issue. Temperature is very near the top of the list of things to monitor, and everyone knows it.
- CBLT 7y agoWould you share what you have in mind? I (personally) would not page in the middle of the night for a temperature issue. The systems should be resilient enough to handle some cpu throttling. A non-transient issue like this would probably end up in a ticket queue.