3 ms·
All it takes is one cluster falling over due to one unexpected issue stemming from unprecedented, more-or-less unreproducible simultaneous load to take a servic
by busted 14y ago
All it takes is one cluster falling over due to one unexpected issue stemming from unprecedented, more-or-less unreproducible simultaneous load to take a service down, especially if it removes the ability to log in.
Distributed services under very heavy load are susceptible to all the same small bugs due to all the normal mistakes every developer makes, except it's much easier for those bugs to cause catastrophic failures!