5 ms·
Digital Ocean completely destroyed my droplet beyond recovery out of the blue
From: DigitalOcean <support@support.digitalocean.com>
Subject: [DigitalOcean Ticket ID:xxxxx] cant power on droplet
Date: August 13, 2014 at 1:26:58 AM GMT-4
To: <xxxxxx@xxxxxxx>
Reply-To: <xxxxxxxx@support.digitalocean.com>
There has been a response to your ticket:
Hello,
You will be unable to power on this droplet as it has suffered unrepairable damage.
We are writing to let you know that the hypervisor that hosts (redacted) has suffered a catastrophic failure. Despite repeated attempts to replace failed components and even to perform two full swaps of the system chassis, we were unable to recover it.
Unfortunately, the failure also resulted in loss of all data on the hypervisor. Droplets and data hosted on this hypervisor node are not recoverable, despite all of our efforts.
Please let us know if you need help recovering from a recent snapshot or backup. While we know that an account credit is only a small comfort when confronted with data loss, we have gone ahead and credited your account for one month of service.
If there is anything else we can do, please let us know. We will be standing by to assist you in recovering from this as best we can, and we have marked the IP that was associated with this Droplet as reserved to your account, so your next Droplet in the SFO datacenter should reclaim it.
Regards,
(redacted)
(edit: line breaks)
- farawayea 12y agoThey are using RAID 5 and probably cheap ssd storage. I'm not surprised. This is what they give you for your money.
- wmf 12y agoSo restore from backup. Hardware failures are inevitable.
- mariust 12y agoI do think that, hardware failures are inevitable, but using a RAID system, should prevent such things since all data is written on 2 or more disks at the same time. The backup might be one or more days old, think of Facebook would say: " hey 1 billion users the 10x billions of images you have uploaded in the last 24h are gone, please upload them again "
- jlawer 12y agoRAID only works when the RAID controller doesn't screw up, or critical metadata doesn't get wiped. If the raid controller starts being the problem your data goes with it. Problem can also happen if your hypervisors filesystem / storage drivers go screwy, get corrupted or one of a million other problems.
- TheLoneWolfling 12y agoRAID is much less effective than it used to be. Current BER (non recoverable read error rate) is around ~1/10^14. If you lose a drive, the probability of you encountering an error when rebuilding the RAID is rapidly approaching 1 as time goes on - BERs are not decreasing nearly as fast as drive sizes are increasing. (This assumes you are using a RAID where you can lose 1 drive and be fine. RAIDs where you can lose 2+ drives are still fine for now.) And this isn't taking into account catastrophic controller failure, damage due to a faulty power supply / surge, etc.
- phaser 12y agoi did. i posted this as a remainder for people to do their backups, especially under digitalocean
- mattkrea 12y agoWhen backups can be enabled at 20% of instance cost I'll deal with it. Nothing beyond development work runs on only one instance due to the problems I've seen even on Amazon.
- Artemis2 12y agoIt's the cloud. You have backups and you can spin up a replacement server in a minute. If you don't have backups, you should probably rethink the way you use these services. Hardware fails, inevitably.
- tuananh 12y agohappened to me once. i wonder what's the actual SSD failure rate at DO? Is it any higher than industry standard rate?
- trampish 12y agoDid you have DO's backup service enabled when this happened? If so, I feel like they should have provided a more streamlined customer experience restoring the damaged/lost droplet from the latest Amazon Glacier instance.
- ksec 12y agoThat is why you should be on Linode is you are in production. Note: Not saying they are perfect, but it is just much less likely to happen.