3 ms·
Yep just had the same email - I've lost all bar one EBS snapshot. Don't think I can trust them anymore, I wish I had more time to move our infrastructure elsew
by pdaddyo 15y ago
Yep just had the same email - I've lost all bar one EBS snapshot. Don't think I can trust them anymore, I wish I had more time to move our infrastructure elsewhere!
They've essentially swiss-cheesed all our backups.
Email copy follows....
Hello,
We've discovered an error in the Amazon EBS software that cleans up unused snapshots. This has affected at least one of your snapshots in the EU-West Region.
During a recent run of this EBS software in the EU-West Region, one or more blocks in a number of EBS snapshots were incorrectly deleted. The root cause was a software error that caused the snapshot references to a subset of blocks to be missed during the reference counting process. This process compares the blocks scheduled for deletion to the blocks referenced in customer snapshots. As a result of the software error, the EBS snapshot management system in the EU-West Region incorrectly thought some of the blocks were no longer being used and deleted them. We've addressed the error in the EBS snapshot system to prevent it from recurring.
We have now disabled all of your snapshots that contain these missing blocks. You can determine which of your snapshots were affected via the AWS Management Console or the DescribeSnapshots API call. The status for any affected snapshots will be shown as "error."
We have created copies of your affected snapshots where we've replaced the missing blocks with empty blocks. You can create a new volume from these snapshot copies and run a recovery tool on it (e.g. a file system recovery tool like fsck); in some cases this may restore normal volume operation. These snapshots can be identified via the snapshot Description field which you can see on the AWS Management Console or via the DescribeSnapshots API call. The Description field contains "Recovery Snapshot snap-xxxx" where snap-xxx is the id of the affected snapshot. Alternately, if you have any older or more recent snapshots that were unaffected, you will be able to create a volume from those snapshots without error. For additional questions, you may open a case in our Support Center: https://aws.amazon.com/support/createCase https://aws.amazon.com/support/createCase
We apologize for any potential impact this might have on your applications.
Sincerely,
AWS Developer Support
- andrewcooke 15y agoare your backups too large to be elsewhere completely? i'm working on a site on appengine, and although i can't back up all data, i can copy the critical stuff (the user accounts) to servers elsewhere. [i realise this is a little off-topic, but what do other app-engine users do?] [edit: wasn't being critical, just trying to understand what others do and why]
- pdaddyo 15y agoThey're not mission critical failures, but that's not the point to me - I pay a monthly fee to have those snapshots there, and they just carved holes in them. Disappointing to say the least, even though it's not actually taken down any of our instances.
- codyrobbins 15y agoI’m glad to hear that you didn’t lose anything critical. But snapshotting EBS volumes are not backing them up. If, by definition, it’s stored using the same service, and therefore susceptible to all the same catastrophes that might befall the service, then it’s not a back up. I always set up a local read slave of the database server which has hot-swappable hard drives that get rotated out of a safe deposit box or fireproof safe. The only way I can properly trust that things are being backed up properly are to do it myself.