3 ms·
Beyond a typing mistake, it's not really very similar. The Gitlab incident was one avoidable problem after another, ending with a giant WTF when they found out
by user15672 10y ago
Beyond a typing mistake, it's not really very similar. The Gitlab incident was one avoidable problem after another, ending with a giant WTF when they found out that no-one had even tested the backups were working.
This is a case of someone slipping on the keyboard, removing more capacity than intended and the recovery process taking longer than expected. The process actually seems to be working (to a given value of working), but the amount of downtime was way above acceptable. They've already put more safeguards into the tooling to prevent the situation from happening again.
S3 is also orders of magnitude more complex than Gitlabs infrastructure, so while the amount of time the outage lasted for is not acceptable, it does show that they at least have working processes for critical situations that allow them to get back in service within a day, which is pretty impressive.